omarsar0· @omarsar0 · X·· 8 小时前AI 评分
omarsar0:Agent 基准都应审计其验证器
omarsar0 表示每个 Agent 基准都应该审计其验证器。他提到 Parsewave 审计了 Zapier 的 AutomationBench 全部 600 个公开任务,让 Agent 编写逼真的错误答案试图欺骗每个验证器,经人工复核确认存在 206 个真实 bug,AutomationBench Verified 已修复全部 206 个。
We need more efforts like this.
Every agent benchmark should audit its verifiers.
Parsewave went through all 600 public tasks in Zapier's AutomationBench.
Agents wrote realistic wrong answers to try to fool each verifier, and human review confirmed 206 real bugs. AutomationBench Verified fixed all 206.
Regarding 1,235 Kimi K3 runs, the fixed verifiers changed 27.9% of the grades.
I just started looking into this benchmark for some independent eval work I am doing, so this is good timing to see this audit.
来源:omarsar0 · x.com