testingcatalog· @testingcatalog · X·· 8 小时前AI 评分
Mistral Large 4 在 DeepSWE 得分 62%,据 VentureBeat 报道超过 GLM-5.3
据 VentureBeat 报道,Mistral Large 4 在 DeepSWE 上得分 62%,超过 GLM-5.3;同时在 Finch 金融任务基准上得分 67%,被称为 SOTA open-weight。此外 Le Chonk 在 Harvey 的 Legal Agent Benchmark 法律任务上得分 15%,同样为 SOTA open-weight。作者表示现在需要一份技术报告,并致谢 @AiBattle_。
据 VentureBeat 报道,Mistral Large 4 在 DeepSWE 上得分 62%,超过了 GLM-5.3。
此外,它在 Finch(金融任务,SOTA 开源权重)上得分 67%。
Le Chonk 在 Harvey 法律智能体基准(法律任务,SOTA 开源权重)上也得分 15%。
我们现在需要一份技术报告 👀
h/t @AiBattle_
来源:testingcatalog · x.com