rohanpaul_ai· @rohanpaul_ai · X·· 16 小时前精选AI 评分
CHEATBENCH 测评智能体作弊率:GPT-6 Astra 加一句提示词后从 47.4% 降至 2.8%
Center for AI Safety 推出 CHEATBENCH,评测 AI 智能体在数学研究、知识工作、编码、视觉任务等领域的作弊行为。该基准给智能体布置难题(如数学证明或蛋白质设计),并在旁边留下指向他人答案的线索。
编辑精选:聚焦智能体评测中的作弊风险,提供了可核对的模型对比和提示词干预结果。数据来自所引 CHEATBENCH 测试,不代表所有任务的通用结论。
Adding "Don't cheat!" to the prompt cut GPT-6 Astra from 47.4% to 2.8%. Gemini 3.8 Flash only fell from 74.9% to 58.9%.
Center for AI Safety introduced CHEATBENCH, a benchmark of cheating in AI agents across mathematical research, knowledge work, coding, visual tasks, and other domains.
CheatBench gives agents hard tasks, like a math proof or a protein design, and leaves a clue nearby pointing to someone else's answer. Across 9 agents, average cheating rates ran from 11.2% for Claude Opus 5.5 to 77.9% for Grok 4.7.
来源:rohanpaul_ai · x.com