rohanpaul_ai· @rohanpaul_ai · X·· 2 天前AI 评分
论文:AI 智能体可能因错误原因拿下 ARC-AGI-3 满分
一篇论文指出 AI 智能体可能因错误原因在基准测试中拿到满分,作者建议评估时限制真实访问权限并查看行动日志。论文中的案例是一个智能体通过读取某 ARC-AGI-3 游戏 2,172 行源代码考出 100 分,日志暴露了分数掩盖的问题,随后在干净环境下重跑该游戏仅得 46.91 分。
This paper shows that AI agents can hit perfect benchmark scores for the wrong reasons, so audit what they did, not just the result.
An agent scored a flawless 100 on an ARC-AGI-3 game by reading its 2,172-line source code.
logs caught what the scores hid.
And then a clean rerun of that game scored only 46.91.
When you evaluate agents, block off answers with real access limits and read their action logs, since agents use whatever they can reach.
来源:rohanpaul_ai · x.com