Artificial Analysis 编码智能体指数新增拒绝时机与回退模型视图
Artificial Analysis 为 Coding Agent Index 新增两项安全拒绝报告视图:拒绝时机(任务提示词即触发还是智能体开始工作后触发)与回退模型(拒绝后切换到哪个模型)。
Safety refusal reporting in the Artificial Analysis Coding Agent Index now shows when refusals occur and which models are used as fallbacks
We've added two ways to explore the results:
➤ Refusal timing: see whether a refusal occurred from the task prompt alone or later in the task, after the agent had already started working
➤ Fallback model: see which model the agent switched to after a refusal, alongside attempts that were blocked
Claude Code with Sonnet 5.5 (max) currently ranks first in the Index. Its safety refusal rate is 4.5%, roughly half the 8.9% recorded for Claude Code with Opus 5.5 (max).
Around 94% of Sonnet 5.5's refusals occurred after the first turn. Following a refusal, the agent almost always switched to Opus 4.8.
Observed fallback patterns vary across model and agent configurations. With Fable 5.1, Claude Code fell back predominantly to Opus 4.8, while Opus 5 accounted for a much larger share of Devin Fusion's fallbacks in the Index.
Both views are available for the overall Index and each benchmark. Safety and fallback behavior is provider-configured and may change over time; these rates reflect behavior recorded at the time of benchmarking.
来源:ArtificialAnlys · x.com