跳到正文
原文
omarsar0· @omarsar0 · X·· 1 天前AI 评分57

Sakana AI 提出 MASS:用多智能体自监督扩展递归自我改进

AI 导读

作者推荐 Sakana AI 关于递归自我改进的论文,提出多智能体自监督方法 MASS。该方法让一个基础模型提出多智能体工作流、运行并打分,用进化搜索保留得分最高的工作流,再用自身轨迹微调模型,改进后的模型进入下一轮循环,从而不再依赖外部验证器。

正文 · 原文

Recommended paper from Sakana AI on recursive self-improvement.

They propose an interesting way to scale recursive self-improvement through multi-agent self-supervision.

In this line of research, self-improvement loops usually need an external verifier, so open-ended tasks without a checker are left out.

MASS removes that requirement.

One base model proposes multi-agent workflows, runs them and grades them, and an evolutionary search keeps the workflows that score best. The model is then fine-tuned on its own traces, and the improved model starts the next cycle as a better optimizer and grader.

Two cycles on Qwen3.6-27B raise performance per output token from 1.2 to 1.6x on four open-ended benchmarks.

A student trained on multi-agent traces also beats a single-agent student trained on 1.4x more tokens.

Paper: arxiv.org/abs/2610.12176

Chat with Paper: academy.dair.ai/papers/recur…

来源:omarsar0 · x.com