跳到正文
原文
omarsar0· @omarsar0 · X·· 2 天前AI 评分52

Google 论文提出 VeriHarness:用分歧与共识双重校验提升智能体答案选择

AI 导读

Google 一篇论文提出 VeriHarness,把同一基础模型变成智能体验证器:一处负责在多次 rollout 出现分歧时核对工作区证据解决争议,另一处挑战所有 rollout 都同意的断言、寻找它们共同遗漏的需求。

正文 · 原文

Banger paper from Google.

It's standard practice to sample several agent rollouts and trust the answers they agree on.

This Google paper shows that agreement can hide shared errors, while disagreement often points to the correct alternative.

VeriHarness turns the same base model into an agentic verifier with two jobs.

One resolves claims where rollouts disagree by checking workspace evidence.

The other challenges claims that every rollout agrees on and looks for requirements they all missed.

Across five long-horizon benchmarks, it gives the best selection scores among the baselines tested. With evidence-backed revision, it adds 6.2 points over a single rollout with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8.

The authors also release about 26,000 rollouts.

Paper: arxiv.org/abs/2610.00972

Chat with Paper: academy.dair.ai/papers/verih…

来源:omarsar0 · x.com