adithya_s_k· @adithya_s_k · X·· 6 天前AI 评分
开源 multi-harness RL 训练方法:在 Claude Code 等 harness 内直接训练模型
作者发布 multi-harness RL 指南,提供一种开放方式,可在 Claude Code、Codex、OpenCode 等实际使用的 agent harness 内部用 RL 训练任意模型于任意任务集,无需改动 harness 或训练代码的任何一行。在四个 harness 上训练后,LFM2.5-2.6B 从 42% 提升到 54%,同时 tool calls 减少 31%。
1/ Excited to release The ultimate guide to multi-harness RL
The same model behaves differently in every agent harness. So we built an open way to train any model with RL on any task set, inside the harnesses people actually use, like Claude Code, Codex, and OpenCode, without changing a single line of harness or training code.
Trained across four harnesses, LFM2.5-2.6B went from 42% to 54% with 31% fewer tool calls. 🧵
来源:adithya_s_k · x.com