Agent 基准需要审计评分器,而不仅是任务本身
Parsewave 审计了 Zapier AutomationBench 的全部 600 个公开任务(前沿实验室在模型卡中引用该基准),用 Agent 生成看似可信的错误答案,再由人工确认哪些评分器失效,结果有 206 个评分器出错,修复后的 AutomationBench Verified 全部修正了这些问题。
Agent benchmarks need audits of the graders, not only the tasks.
Parsewave reviewed all 600 public tasks in Zapier's AutomationBench, which frontier labs cite on their model cards.
It used agents to write convincing wrong answers, then had humans confirm which graders actually failed: 206 did.
AutomationBench Verified fixes every one.
On 1,235 Kimi K3 runs, the fixed graders gave a different verdict 27.9% of the time.
The clearest case is task 813. The grader checked the Salesforce notes but never looked at the DocuSign template. A submission that sent four contracts on the Standard template, instead of the required GDPR, HIPAA, SOC2 and Enterprise ones, passed 13 of 13 checks and scored 1.0. After the fix it scores 0.09.
来源:rohanpaul_ai · x.com