跳到正文
灵耀AI热点 LEADERBOARD

AI 模型排行榜

数学、逻辑与陌生规则,看模型能不能想明白新问题。

5 项评测·2 家机构·10/03 14:10 更新

共识指数 TOP 12

纵轴从 77 起,只放大差距,不代表 0 分;点柱进详情。

  1. 01gpt-6.1-sol
  2. 02GPT-6 Astra
  3. 03Claude Opus 5.5
  4. 04Claude Fable 5.1
  5. 05GPT-5.6 Sol
  6. 06claude-sonnet-5.5
  7. 07Claude Opus 5
  8. 08Claude Fable 5
  9. 09GPT-6 Sol
  10. 10GPT-5.6 Terra
  11. 11GPT-5.5
  12. 12Claude Opus 4.8

推理榜的名次来自共同证据,柱高按本榜实际分差等比:名次与分数偶有不同序,属方法本身的取舍。

价格 vs 能力

  • OpenAI
  • Anthropic
  • Meta
  • Google
  • xAI
  • DeepSeek
  • Alibaba
  • Moonshot AI
  • Z.ai

越靠右上越划算(便宜且分高)。横轴对数刻度,右端最便宜;点进模型详情。

越划算 ↗5060708090100¥10¥100横轴 输出价格 · 人民币 / 百万 Token(对数,右端最便宜)

28 个模型有官方输出价格,另有 2 个未核验价格未进图(gpt-6.1-sol、claude-sonnet-5.5)。价格不参与排名,只用来读这张图;名字放不下时不画标签,悬停或点开仍可看到完整信息。

GPT-6 Astra(OpenAI)输出 ¥335,共识指数 98.5,第 2 名;Claude Opus 5.5(Anthropic)输出 ¥134,共识指数 97.2,第 3 名;Claude Fable 5.1(Anthropic)输出 ¥335,共识指数 95.5,第 4 名;GPT-5.6 Sol(OpenAI)输出 ¥134,共识指数 94.2,第 5 名;Claude Opus 5(Anthropic)输出 ¥168,共识指数 89.5,第 7 名;Claude Fable 5(Anthropic)输出 ¥335,共识指数 88.3,第 8 名;GPT-6 Sol(OpenAI)输出 ¥67,共识指数 87.7,第 9 名;GPT-5.6 Terra(OpenAI)输出 ¥80.5,共识指数 87.1,第 10 名;GPT-5.5(OpenAI)输出 ¥201,共识指数 84.4,第 11 名;Claude Opus 4.8(Anthropic)输出 ¥168,共识指数 79.4,第 12 名;Muse Spark 1.3(Meta)输出 ¥28.5,共识指数 79.3,第 13 名;Gemini 3.7 Flash(Google)输出 ¥25.1,共识指数 76.5,第 14 名;Grok 4.6(xAI)输出 ¥40.2,共识指数 75.5,第 15 名;DeepSeek V4 Pro 0813(DeepSeek)输出 ¥27,共识指数 74.1,第 16 名;Gemini 3.8 Flash(Google)输出 ¥25.1,共识指数 74.1,第 17 名;GPT-5.6 Luna(OpenAI)输出 ¥8.05,共识指数 72.3,第 18 名;Qwen3.8 Max(Alibaba)输出 ¥36,共识指数 71.8,第 19 名;Claude Sonnet 5(Anthropic)输出 ¥67,共识指数 70.8,第 20 名;Claude Opus 4.7(Anthropic)输出 ¥168,共识指数 67.9,第 21 名;GPT-5.2 2025-12-11(OpenAI)输出 ¥93.9,共识指数 63.9,第 22 名;Kimi K3(Moonshot AI)输出 ¥100,共识指数 62.4,第 23 名;Grok 4.7(xAI)输出 ¥40.2,共识指数 62.2,第 24 名;Grok 4.5(xAI)输出 ¥40.2,共识指数 61.0,第 25 名;Claude Opus 4.6(Anthropic)输出 ¥168,共识指数 59.2,第 26 名;Gemini 3.1 Pro Preview(Google)输出 ¥80.5,共识指数 57.1,第 27 名;GLM-5.3(Z.ai)输出 ¥29.5,共识指数 55.4,第 28 名;DeepSeek V4 Flash 0731(DeepSeek)输出 ¥9,共识指数 50.1,第 29 名;Gemini 3.6 Flash(Google)输出 ¥25.1,共识指数 47.7,第 30 名

推理榜TOP 30

当前展示 30 个模型,可按列重排;名次来自原榜
模型
01gpt-6.1-solopenai上线 — · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 —输入 — · 输出 —98.7
02GPT-6 AstraOpenAI上线 2026-09-03 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥6.7输入 ¥67.05 · 输出 ¥335.2398.5
03Claude Opus 5.5Anthropic上线 2026-09-17 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Mystery Game Puzzles缓存输入 ¥1.34输入 ¥26.82 · 输出 ¥134.0997.2
04Claude Fable 5.1Anthropic上线 2026-09-01 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥1.68输入 ¥67.05 · 输出 ¥335.2395.5
05GPT-5.6 SolOpenAI上线 2026-07-09 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥2.68输入 ¥26.82 · 输出 ¥134.0994.2
06claude-sonnet-5.5anthropic上线 — · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Mystery Game Puzzles缓存输入 —输入 — · 输出 —92.8
07Claude Opus 5Anthropic上线 2026-07-24 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥3.35输入 ¥33.52 · 输出 ¥167.6289.5
08Claude Fable 5Anthropic上线 2026-06-09 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥6.7输入 ¥67.05 · 输出 ¥335.2388.3
09GPT-6 SolOpenAI上线 2026-09-22 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Mystery Game Puzzles缓存输入 ¥1.34输入 ¥13.41 · 输出 ¥67.0587.7
10GPT-5.6 TerraOpenAI上线 2026-07-09 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥1.34输入 ¥13.41 · 输出 ¥80.4687.1
11GPT-5.5OpenAI上线 2026-04-23 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥3.35输入 ¥33.52 · 输出 ¥201.1484.4
12Claude Opus 4.8Anthropic上线 2026-05-28 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥3.35输入 ¥33.52 · 输出 ¥167.6279.4
13Muse Spark 1.3Meta上线 2026-09-02 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 —输入 ¥8.38 · 输出 ¥28.4979.3
14Gemini 3.7 FlashGoogle上线 2026-08-13 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥0.5输入 ¥5.03 · 输出 ¥25.1476.5
15Grok 4.6xAI上线 2026-08-12 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥3.35输入 ¥13.41 · 输出 ¥40.2375.5
16DeepSeek V4 Pro 0813DeepSeek上线 2026-08-13 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥0.3输入 ¥9 · 输出 ¥27官方权重 ↗74.1
17Gemini 3.8 FlashGoogle上线 2026-09-02 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥0.5输入 ¥5.03 · 输出 ¥25.1474.1
18GPT-5.6 LunaOpenAI上线 2026-07-09 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥0.13输入 ¥1.34 · 输出 ¥8.0572.3
19Qwen3.8 MaxAlibaba上线 2026-07-19 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 —输入 ¥12 · 输出 ¥3671.8
20Claude Sonnet 5Anthropic上线 2026-06-30 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥1.34输入 ¥13.41 · 输出 ¥67.0570.8
21Claude Opus 4.7Anthropic上线 2026-04-16 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥3.35输入 ¥33.52 · 输出 ¥167.6267.9
22GPT-5.2 2025-12-11OpenAI上线 2025-12-11 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥1.17输入 ¥11.73 · 输出 ¥93.8663.9
23Kimi K3Moonshot AI上线 2026-07-16 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥2输入 ¥20 · 输出 ¥100官方权重 ↗62.4
24Grok 4.7xAI上线 2026-09-21 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥3.35输入 ¥13.41 · 输出 ¥40.2362.2
25Grok 4.5xAI上线 2026-07-08 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles缓存输入 ¥2.01输入 ¥13.41 · 输出 ¥40.2361.0
26Claude Opus 4.6Anthropic上线 2026-02-05 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥3.35输入 ¥33.52 · 输出 ¥167.6259.2
27Gemini 3.1 Pro PreviewGoogle上线 2026-02-19 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥1.34输入 ¥13.41 · 输出 ¥80.4657.1
28GLM-5.3Z.ai上线 2026-08-18 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥1.74输入 ¥9.39 · 输出 ¥29.5官方权重 ↗55.4
29DeepSeek V4 Flash 0731DeepSeek上线 2026-07-31 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥0.1输入 ¥3 · 输出 ¥9官方权重 ↗50.1
30Gemini 3.6 FlashGoogle上线 2026-07-21 · 证据敏感评测来源:LiveBench · 推理与数学、FrontierMath v2 · Tiers 1–3、FrontierMath v2 · Tier 4、Chess Puzzles、Mystery Game Puzzles缓存输入 ¥0.5输入 ¥5.03 · 输出 ¥25.1447.7

按多项公开评测的共同证据排名。每个榜单或筛选结果最多展示 30 个模型,筛选后保留原榜名次与分数。

共识指数不是正确率;同分仍按共同证据确定的名次展示。

如何看这张榜

综合多家公开评测,不同模型的参评覆盖不同。价格不参与排名,缺测不记零分,指数不是正确率。

了解计算方法 →

关于价格

API 价格来自厂商官网,按每百万 Token 展示。美元报价按 2026-10-02 汇率折算成人民币。缓存价格指命中后的输入价格,缓存写入、存储及订阅费用另计。