跳到正文
原文
omarsar0· @omarsar0 · X·· 20 小时前AI 评分59

SelfSearch 让编码智能体自改 harness,以 $4.03 成本在 Terminal-Bench 2.1 追平 Codex

AI 导读

一篇 arXiv 论文提出 SelfSearch,让编码智能体反复改写自己的 harness(指令、工具与流程),用 DeepSeek V4 Flash 在 Terminal-Bench 2.1 上达到 82.0%,搜索成本仅 $4.03,在同一设置下的九款 harness 公开对比中追平顶级的 Codex。

正文 · 原文

Learn to optimize your own harness, folks.

You can squeeze much more performance and value from a harness.

This work claims that a self-modified harness matched Codex for $4.03.

Specifically, a coding agent rewrote its own harness until it solved 82.0% of Terminal-Bench 2.1 with DeepSeek V4 Flash, without any task reward during the search.

The search cost was $4.03.

That matches Codex, the top harness in a public nine-harness comparison run under the same settings.

SelfSearch has agents modify their own instructions, tools, and procedures using records of earlier self-modification attempts. Each record holds the reasoning, tool actions, and outcomes. The modified agent then becomes the next improver.

Population-mean success rises in all six model-benchmark settings, with single agents gaining up to 11.2 points. On SWE-bench Multilingual, one evolved agent gains 5.0 points and spends 38.5% less on tasks both versions solve.

Paper: arxiv.org/abs/2609.37968

Chat with Paper: academy.dair.ai/papers/selfs…

来源:omarsar0 · x.com