X 用户讨论 Artificial Analysis 成本图表为何使用对数刻度
X 上一段动画直观演示了 Artificial Analysis 成本图表为何采用对数刻度——Cost per Task 的 X 轴覆盖 1000x 的价格区间,线性刻度无法呈现。评论区还对比了性价比:Muse Glimmer 约 340 分/$1,Claude Opus 5.5 约 16 分/$1,Claude Fable 5.1 约 9.5 分/$1,但高分模型解锁的能力不可替代。
Oh, thaaat's why they use a logarithmic scale!
![]()
Oct 3, 2026 · 8:19 PM UTC
16
43 1,986 162,657
2
3 3,228
True, plus focusing on Max makes it look bad. Their max is like 2-3x as expensive as xhigh.
2
2,335
Cool animation! This is why this chart is so important - the X axis is showing you a 1000x range in Cost per Task. Maybe we should add a button to show this on the homepage!
1
3 414
You could, but it makes the chart kind of unreadable, tbh Anyway, I love Artificial Analysis. I was wondering if you could take a look at my side project, mypriors.com, which is kind of like Artificial Analysis, but you get to customize all the benchmarks that it shows and what mix you use. So if you disagree with the rankings, you can change them. I would love any feedback you have!
262
massive diminishing returns in intelligence/$ Muse Glimmer ∼17 / $0.05 = ∼340 points per $1 Claude Opus 5.5 ∼56 / $3.5 = ∼16 per $1 Claude Fable 5.1 ∼57 / $6 = ∼9.5 per $1 but those final intel points are between 'getting it or not', so it's still worth paying for
3
4 3,291
Points per dollar is not really a good metric, though, because the points are not really interchangeable. A model that scores in the 50s is in an entirely different league of intelligence than Muse Glimmer. In the real world, the usefulness it can give you isn't just like three times Muse Glimmer if it's performing thrice as well. It's an entirely different kind of capability that unlocks things Muse Glimmer could never do, no matter how many tokens you spend on it. Not to mention the value of reliability. That means instead of having to babysit one thread to make sure it's not doing dumb stuff, you can send off many threads at a time, some of which you don't even need to check because you can trust that they'll deliver the outcome you specified.
7
1,382
This is the first time I actually learnt something on X
1
13 2,497
Glad to be of assistance 🫡
3
2,162
Interesting that the cost is in log scale but the Reasoning score is not.
1
3 1,213
The value of a log scale is when your data is covering so many orders of magnitude that it would be impractical to show it on a linear scale. The capability score doesn't. Intelligence scores are ranging basically from 10 to 60 right now, not even one order of magnitude. A logarithmic scale would kind of just be silly.
3
1,002
Taps the sign:
Dare I say, this is the correct version of the chart.
1
230
Not a single thing has been made or done yet with a Anthropic model that was worth the higher price, even when accounting for that people pay for the competitive edge they allegedly get with it.
299
man if u just show this simple animation in math class for explaining logarithm, u would be able to one shot even the lowest iq one.... such a great way to visualize it 👌
1
62 3,506
The point of log scale is to show proportionality that is important. The difference in .001 and .01 per mTok in a linear scale are imperceptible, but the cost of usage in practice is. You're exaggerating the difference by using a small interval
ALT A more reasonable scale for token costs, albeit log scales make more sense
ALT A ridiculous, impractical scale to demonstrate token cost
1
13 2,207
The most misleading of scales!
2
1,629
This is such a cool and intuitive way of visualising the logarithmic scale.
19
2,769
This is violence.
1
1,503
people learning about scaling laws
83
Claude models are cheap on subscription and very costly on API. Well not so cheap but usable since 5.5 release.
1
421
WHAT THE FUCK
3
1,059
OOOOH
1
65
来源:X:Peter Steinberger (@steipete) · x.com


















