Nvidia 论文提出 VERA:交替训练模型与编辑技能文件让长流程 Agent 提升最大
Nvidia 新论文提出 VERA,让长多步任务的 Agent 交替进行模型训练与技能文件编辑,并用真实证据的逐步打分决定每次修改方向。论文称只训练模型或只改 harness 会损失约一半收益;VERA 将 benchmark 运行转成 9,000 多个可重启沙盒,逐步对照真实文件和日志打分。
New Nvidia paper shows agents for long, multi-step work improve most when you alternate between training the model and editing its skill files, using step-by-step scores from real evidence to pick each fix.
Training only the model or only the harness leaves about half the gain on the table, compared with updating both in alternating rounds.
Most environments score only the final result, which hides which step broke. VERA builds over 9,000 restartable sandboxes that check each step against real files and logs, then improves the agent in rounds.
VERA turns benchmark runs into over 9,000 restartable sandboxes where each step of a long workflow gets its own checklist score from real evidence.
On a medical research benchmark, a 9B agent scored 69.1 with both kinds of updates, versus 56.1 with skill edits alone and 43.3 with training alone.
Score each step against real artifacts, and let those scores decide whether the next fix goes into the model or its skills.
来源:rohanpaul_ai · x.com