跳到正文
原文
omarsar0· @omarsar0 · X·· 17 小时前AI 评分59

作者呼吁所有智能体基准都应审计其验证器

AI 导读

作者转发了 Parsave 对 Zapier AutomationBench 的审计工作:其检查了全部 600 个公开任务,用逼真的错误答案试图欺骗各验证器,经人工复核确认 206 个真实缺陷,AutomationBench Verified 已修复全部 206 个。

正文 · 原文

We need more efforts like this.

Every agent benchmark should audit its verifiers.

Parsave went through all 600 public tasks in Zapier's AutomationBench.

Agents wrote realistic wrong answers to try to fool each verifier, and human review confirmed 206 real bugs. AutomationBench Verified fixed all 206.

Regrading 1,235 Kimi K3 runs with the fixed verifiers changed 27.9% of the grades.

I just started looking into this benchmark for some independent eval work I am doing, so this is good timing to see this audit.

来源:omarsar0 · x.com