omarsar0· @omarsar0 · X·· 1 天前AI 评分
RSIGym 用系统工程提升自我改进智能体
RSIGym 为研究型智能体提供训练、推理、评测和沙箱即服务,使其把预算花在实验而非重建基础设施上;以 Opus 5 作为研究者,改进后的系统在 SWE-bench Verified 上从 17.67% 提升到 50.33%。推文还提到一种衡量 harness 与模型协同进化质量的方法,这正是全栈 AI 公司保持前沿的方式。
Recommended read. And I agree that the RSI is also a systems engineering problem.
Self-improving agents need better research environments.
RSIGym gives a research agent training, inference, evals, and sandboxes as services it can call. The agent spends its budget on experiments instead of rebuilding infra.
With Opus 5 as the researcher, the improved system went from 17.67% to 50.33% on SWE-bench Verified.
Also cool to see a way to measure the quality of co-evolution between harnesses and models, which is how full-stack AI companies stay on the frontier.
来源:omarsar0 · x.com