ActiveSaddler 动态调整智能体训练场景提升性能
微软等团队提出 ActiveSaddler,通过将反复出现的失败归类为失败模式并作为非平稳 bandit 的 arm,动态调整智能体 harness 的训练场景。该方法追踪各失败模式的学习收益,在复访已知弱点与发现新问题间分配预算,在相同优化器下,GAIA2 Pass@1 提升 4.4 点,Terminal-Bench 2.0 提升 7.5 点。
Great paper from Microsoft and colleagues on optimizing agent harnesses.
Current harness optimizers change how the harness is updated but keep the training scenarios fixed, so feedback keeps coming from tasks that stop being informative as the harness improves.
This work adapts the scenarios as well.
ActiveSaddler groups recurring failures into failure patterns and treats each pattern as an arm in a non-stationary bandit.
It tracks how much the harness is still learning from each pattern and splits the budget between revisiting known weaknesses and finding new ones.
With the same optimizer, test Pass@1 improves by 4.4 points on GAIA2 and 7.5 points on Terminal-Bench 2.0 compared with a fixed scenario order.
Paper: academy.dair.ai/papers/activ…
来源:dair_ai · x.com