Claude Opus 5 作为自动化 AI 研究员将 Qwen 模型 SWE-bench Verified 分数提升近三倍
Evolvent_AI 新发布的开源 RSIGym 让 AI Agent 在同一预算环境中重训模型并重写 harness,使 Claude Opus 5 作为自动化 AI 研究员把一个 Qwen 模型的 SWE-bench Verified 分数提升近三倍。
Claude Opus 5 nearly tripled a Qwen model's SWE-bench Verified score while working as an automated AI researcher.
@Evolvent_AI 's newly released open-source RSIGym made that measurement possible by letting an AI agent retrain a model and rewrite its harness inside one budgeted environment.
shows that frontier AI agents can substantially improve another AI model when training, serving, and testing come as ready-made services.
Agents work from CPU-only containers and call remote services for LoRA fine-tuning, model serving, benchmarking and sandboxes, all charged against a per-run budget.
In the main test, 6 frontier agents started from Qwen3.5-35B-A3B-Base and a minimal harness, with $500 of services per benchmark.
If you're improving an agent, read its failure logs and fix the harness first: heavier training scored lower in 8 of 10 small tests.
来源:rohanpaul_ai · x.com