CMU 论文提出用 RL 训练模型改写 agent harness 代码
DAIR.AI 转发 CMU 关于 harness learning 的论文,作者用 RL 训练一个 proposer 模型,读取任务、当前 harness 和执行报告后直接写出 harness 代码修改,奖励是修改后 harness 的得分,solver 模型保持不变。
Banger paper from CMU on harness learning.
(bookmark it)
Also, pay attention to this important new AI engineering skill of improving agents by editing their harness code instead of their weights.
Seeing a huge shift towards this.
The authors train a proposer model with RL to read a task, the current harness and an execution report, then write a code edit to the harness.
The reward is the score of the revised harness. The solver model never changes.
A trained 4B proposer beats its 35B teacher at single-step revision on Reasoning Gym, including task families it never saw in training. A proposer trained on HotpotQA keeps improving harnesses on MuSiQue and 2WikiMultihopQA.
Paper: arxiv.org/abs/2609.35738
Chat with Paper: academy.dair.ai/papers/harne…
来源:omarsar0 · x.com