跳到正文
原文
dair_ai· @dair_ai · X·· 2 天前AI 评分62

NVIDIA 论文提出 Mid-Harness:终端智能体用验证器提升测试时算力效率

AI 导读

NVIDIA 一篇关于终端智能体测试时算力的论文提出 Mid-Harness 方法:采样多条候选 shell 命令并先验证再执行,把更多算力花在验证器而非增加采样上。

正文 · 原文

Banger paper from NVIDIA on test-time compute for terminal agents.

The finding is that you should sample several candidate shell commands, verify them before running one, and spend more on the verifier than on extra samples.

With a GPT-5.6 Sol verifier choosing among 8 sampled actions, TerminalBench-Lite Pass@1 rises from 50.0% to 68.0%. With a weak verifier, extra samples add almost nothing.

Mid-Harness leaves the generator and harness unchanged and works between them. When a small TMAX-9B model verifies its own candidates, pairwise comparison works best, and distilling the strong verifier into it helps further.

Combining action sampling with trajectory sampling reaches higher success at lower estimated token cost than sampling full trajectories alone.

Paper: academy.dair.ai/papers/mid-h…

来源:dair_ai · x.com