跳到正文
原文
X:Karina Nguyen(@karinanguyen)· @ShayanCJahan·· 2 天前AI 评分41

Repovive 数学竞赛 AI 判卷评测:GPT 与 Claude 模型评判数学证明的能力对比

RT by @karinanguyen: 1/4 How well can models judge mathematical reasoning? After our first math contest on Repovive, we worked with a team of IOI, IMO, and ICPC finalists to evaluate our AI judge. We selected 534 suspicious or borderline submissions for review, established a shared grading standard, and created expert reference judgments. We then compared models under different grading instructions against those judgments. The main question: can models reliably understand and check challenging mathematical proofs? Accuracy is higher across the full set of contest submissions.

AI 导读

Repovive 联合 IOI、IMO、ICPC 决赛选手评测 AI 判卷能力,从数学竞赛中选出 534 份可疑或边缘提交,建立统一评分标准与专家参考判断。模型整体偏严,更常拒绝有效提交;清晰指令让 GPT 模型平均提升约 2 个百分点,加入评分示例后 Claude 平均提升约 10 个百分点、GPT 约 5 个百分点。Grok 几乎拒绝一切,仅接受 1–5% 提交,而金标准接受 64.4%。

来源:X:Karina Nguyen(@karinanguyen) · x.lingyaoai.com