跳到正文
原文
omarsar0· @omarsar0 · X·· 13 小时前AI 评分62

论文:HERMES harness 将 GPT-5.6 Sol 仓库迁移成功率从 6.5% 提升至 31.0%

AI 导读

作者引述一篇论文指出,在整仓库迁移任务中,同一 GPT-5.6 Sol 模型与相同 effort 设置下,把 Codex 换成 HERMES harness 后成功率从 6.5% 升至 31.0%。

正文 · 原文

Build your own harness, folks.

Reading papers like this makes me realize how underexplored harness engineering really is.

The authors find that on whole-repository migration, GPT-5.6 Sol goes from 6.5% to 31.0% when Codex is replaced with the HERMES harness, with the same model and effort setting.

The gain comes from the harness.

HERMES pairs each repository component with a resident LLM that knows its own code and dependencies. A dependency-aware step decides which components to activate, and a diagnosis step maps test failures back to the components that need changes.

Across four software engineering benchmarks, it beats matched baseline harnesses by 12.4 points on average.

With strong activation and diagnosis models, Qwen3-8B components come within 4.5 points of an all-GPT-5.6 Sol setup and cut Terminal-Bench 4.0 inference cost by 26.2%.

Paper: arxiv.org/abs/2610.07832

Chat with Paper: academy.dair.ai/papers/harne…

来源:omarsar0 · x.com