跳到正文
原文
pwendell· @pwendell · X·· 7 天前AI 评分74

Databricks 测评:Opus 5.5 与 GPT-6 Luna 明显推进成本质量前沿

AI 导读

Databricks 的 pwendell 发布基于 N=2,400 名工程师在线负载分析与离线评测的结果,认为上周发布的三款模型中有两款明显扩展成本质量前沿:Opus 5.5 和 GPT-6 Luna。

正文 · 原文

Crazy few weeks for model releases! Our findings @Databricks show several new models meaningfully advance the pareto frontier. Results below (online workload analysis of N=2,400 engineers, plus offline evals):

1. Two of the three models released last week clearly expand the cost/quality frontier: Opus 5.5 and GPT-6 Luna.

2. Opus 5.5 is now the highest quality mid-tier model. It is better than all prior Opus models, better than GPT-6 Sol, and better than GPT-5.6 Sol.

3. Opus 5.5 reduces same-task costs consistently by 20% in both offline and online analysis. This is against a baseline of Opus 4.8, the prior least-cost Opus model (Opus 5.0 was a bit of a dud with high costs and barely noticeable quality improvements).

4. Due to best-in-class quality and lower costs, Opus 5.5 is a strong candidate as an “every day default” model for coding, and we are now encouraging it for this purpose at Databricks.

5. GPT-6 Luna is very, very, very cheap. It was at least 20 times cheaper per-task than Opus 5.5 in every offline benchmark we tested and in observed online use.

6. GPT-6 Luna is surprisingly capable given how cheap it is. On one of our most difficult evaluation suites it roughly matches Opus 4.6 performance, while being 99.3% cheaper per-task than Opus 4.6 was at that time. That's a 100X cost reduction in ~9 months! This finding is preliminary and we are still evaluating Luna quality on a broader set of offline and online tests.

Our production setup: Unity Gateway to route workloads across models and trace agentic interactions. A mix of end-user harnesses including: Omingent (meta-harness), Claude Code, Codex, and Cursor.

来源:pwendell · x.com