跳到正文
原文
X:Arena (@arena)· @arena·· 2 天前精选AI 评分78

Arena 评测:Gemini 4 Argon (High) 登顶 Text Arena,Agent Arena 排第 8

RT by @arena: More big news from @GoogleDeepMind: Gemini 4 Argon (High) is #8 in Agent Arena with a +7.92% net improvement score, and has reshaped the Pareto frontier with a $0.62 cost per task! See its placement below. Gemini 4 Argon (High) is a 4.96 percentage point improvement over Gemini 3.8 Flash (High), at #19 with +2.96% net improvement. By key signals, Gemini 4 Argon (High) stands out in: - #1 in Steerability with +15.88% (the model’s ability to course-correct when you push back) - #2 in Confirmed Success with +14.15% (explicit user feedback that the task worked) - #4 in Praise vs Complaint with +27.74% (implicit sentiment in user reactions) By category Gemini 4 Argon (High) is especially strong in Chat, landing at #3 with +11.58% net improvement. With 3k real-world agentic sessions so far, this score is preliminary. Stay tuned as more traces come in from our global community of users. Congrats to the @GoogleDeepMind team on this release!

AI 导读

Arena 公布 Google DeepMind 的 Gemini 4 Argon (High) 在 Text Arena 以 1525 分排名第 1,在 Code Arena: WebDev 以 1679 分排名第 8,在 Agent Arena 以 +7.92% 净提升排名第 8,每任务成本 0.62 美元。

推荐理由

Arena 公布了 Gemini 4 Argon (High) 在 Agent Arena 与 Text Arena 的排名和成本数据,可据此比较其与 Gemini 3.8 Flash、Claude Opus 4.6 的差距。

来源:X:Arena (@arena) · x.lingyaoai.com