testingcatalog· @testingcatalog · X·· 10 小时前AI 评分
Mistral Large 4 在 DeepSWE 得分 62%,据 VentureBeat 报道超过 GLM-5.3
据 VentureBeat 报道,Mistral Large 4 在 DeepSWE 上得分 62%,超过 GLM-5.3;同时在 Finch 金融任务基准上得分 67%,被称为 SOTA open-weight。此外 Le Chonk 在 Harvey 的 Legal Agent Benchmark 法律任务上得分 15%,同样为 SOTA open-weight。作者表示现在需要一份技术报告,并致谢 @AiBattle_。
Mistral Large 4 scores 62% on DeepSWE, outperforming GLM-5.3, according to VentureBeat.
Additionally, it scores 67% on Finch (Financial tasks, SOTA open-weight).
Le Chonk also scores 15% on Harvey’s Legal Agent Benchmark (Legal tasks, SOTA open-weight).
We need a tech report now 👀
h/t @AiBattle_
来源:testingcatalog · x.com