Meta Superintelligence Labs 论文提出 agent plasticity 度量自我改进智能体
Meta Superintelligence Labs 发表关于自我改进智能体的论文,提出 agent plasticity 指标,即在模型权重冻结、每次运行从全新上下文开始的条件下,每投入 1 美元学习成本在保留任务上的增益。
Banger paper from Meta Superintelligence Labs on self-improving agents.
(bookmark it)
It's hard to know exactly what drives self-improvement, since so many variables are at play (notes, skills, tool calls between runs, etc.).
Meta researchers explore and discuss a way to measure whether self-improvement pays off.
They call it agent plasticity. It is the gain on held-out tasks per dollar spent on learning, with model weights frozen and every run starting from a fresh context.
They find that the model that performs the best is often a different model from the one that learns most efficiently.
In chess, Go, and Hex, Claude Fable 5 reaches the highest final score, while GPT-5.6 Sol gains the most per dollar.
In NetHack, only Claude Opus 5.5 improves significantly, by 66 normalized points for about $1,073 of learning.
Another interesting finding is that slow learners often ignore artifacts they already wrote. Faster learners reuse their artifacts and still fail when an artifact is low quality.
Paper: arxiv.org/abs/2610.08902
Chat with Paper: academy.dair.ai/papers/agent…
来源:omarsar0 · x.com