SelfSearch 让编码智能体自改 harness,以 $4.03 成本在 Terminal-Bench 2.1 追平 Codex
一篇 arXiv 论文提出 SelfSearch,让编码智能体反复改写自己的 harness(指令、工具与流程),用 DeepSeek V4 Flash 在 Terminal-Bench 2.1 上达到 82.0%,搜索成本仅 $4.03,在同一设置下的九款 harness 公开对比中追平顶级的 Codex。
Learn to optimize your own harness, folks.
You can squeeze much more performance and value from a harness.
This work claims that a self-modified harness matched Codex for $4.03.
Specifically, a coding agent rewrote its own harness until it solved 82.0% of Terminal-Bench 2.1 with DeepSeek V4 Flash, without any task reward during the search.
The search cost was $4.03.
That matches Codex, the top harness in a public nine-harness comparison run under the same settings.
SelfSearch has agents modify their own instructions, tools, and procedures using records of earlier self-modification attempts. Each record holds the reasoning, tool actions, and outcomes. The modified agent then becomes the next improver.
Population-mean success rises in all six model-benchmark settings, with single agents gaining up to 11.2 points. On SWE-bench Multilingual, one evolved agent gains 5.0 points and spends 38.5% less on tasks both versions solve.
Paper: arxiv.org/abs/2609.37968
Chat with Paper: academy.dair.ai/papers/selfs…
来源:omarsar0 · x.com