arcprize· @arcprize · X·· 8 小时前AI 评分
ARC Prize 分析 Grok 4.7 reasoning token 用量与 ARC-AGI-2 得分的关联
ARC Prize 指出 Grok 4.7 的 reasoning token 用量与其 ARC-AGI-2 公开得分相关:low 档每次 test-pair 尝试平均约 10k tokens、得分 25%,medium 至 xhigh 档则为 86k 至 120k tokens、得分 57.5% 至 60%。作者认为 low 档较低的 token 用量或可解释其较低的得分。
Grok 4.7's reasoning-token usage correlates with its ARC-AGI-2 public score: low averaged 10k tokens per test-pair attempt and scored 25%, versus 86k to 120k tokens and 57.5% to 60% at medium through xhigh. Low's lower token usage may help explain its lower score.
来源:arcprize · x.com