ProVer:用 LLM 裁判定位关键步骤,改进智能体 RL 的信用分配
Good paper on credit assignment for agent RL. The main finding is that you want an LLM judge to choose where to check a trajectory, and the rollouts to decide how much credit that step gets. GRPO gives every token in a trajectory the same advantage, so the training signal cannot tell the decisive step from the rest. ProVer has a judge compare successful and failed rollouts and name the segment it thinks caused the difference. It then samples continuations from just before and just after that segment and uses the change in success rate as the segment's advantage. Across ALFWorld, WebShop and SearchQA, this gives relative improvements over GRPO of 9.91% for Qwen3.5-2B and 7.12% for Qwen3.5-4B. It still helps when the judge is a smaller model. Paper: https://arxiv.org/abs/2609.36178 Chat with Paper: https://academy.dair.ai/papers/targeting-pivotal-decisions-for-credit-assignment-in-agentic-reinforcement-learn-2609.36178
ProVer 提出一种智能体强化学习信用分配方法:由 LLM 裁判对比成功与失败的 rollout,指出造成差异的关键片段,再通过该片段前后采样续写、以成功率变化作为该片段的优势值。在 ALFWorld、WebShop 和 SearchQA 上,相对 GRPO 分别带来 Qwen3.5-2B 9.91%、Qwen3.5-4B 7.12% 的相对提升,裁判换成更小模型时仍有效。
来源:X:Elvis Saravia (@omarsar0, DAIR.AI) · x.lingyaoai.com