面壁智能发布 JustRL II:为 GRPO 加入 critic 做 token 级信用分配
RT by @OpenBMB: Introducing JustRL II 🚀 Building on JustRL (https://x.com/HBX_hbx/status/1988474153436090776), we took a closer look at how GRPO behaves in long-CoT RL (128k). The group-mean baseline is a great fit for short traces, but over tens of thousands of tokens it gives a coarse, response-level signal. JustRL II keeps GRPO's group structure and adds a critic for token-level credit assignment. Just as simple, keeps improving where GRPO levels off: AIME25 61→81 on a 2B model. The same recipe powers the RL stage of MiniCPM5-2B (https://huggingface.co/openbmb/MiniCPM5-2B), making it SOTA among models under 4B. Data + checkpoints are open. Code lands this week. 📖 https://panhaoxuan.notion.site/justrl-ii-scaling-small-llms-to-128k-reasoning-with-a-critic
面壁智能发布 JustRL II,在保留 GRPO 分组结构的基础上加入 critic 做 token 级信用分配,以改善长 CoT RL(128k)中组均值基线只给出粗粒度响应级信号的问题。该方法在 2B 模型上把 AIME25 从 61 提升到 81,并用于 MiniCPM5-2B 的 RL 阶段,使其在 4B 以下模型中达到 SOTA。数据与 checkpoint 已开放,代码将于本周发布。
来源:X:面壁智能 OpenBMB (@OpenBMB) · x.lingyaoai.com