面壁智能 OpenBMB 推出 Diffusion Reward Models(DRM):学习完整奖励分布而非单一分数
A reward of “3” can mean two completely different things. Everyone thinks a response is mediocre — or half the people love it while the other half hate it. Most Reward Models cannot tell the difference. Introducing Diffusion Reward Models (DRM): instead of collapsing human preference into a single score, DRM learns the full reward distribution, preserving disagreement, uncertainty, and multiple plausible judgments. ✨ Paper:https://arxiv.org/abs/2609.33803 🤗 Models: https://huggingface.co/Teburile/DRM 💻 GitHub: https://github.com/thunlp/DRM Why it matters: Human disagreement is structured, not just noise. On datasets with repeated annotations, judgments often form separated or polarized patterns. More importantly, as human disagreement increases, DRM’s learned reward distribution becomes increasingly multimodal. The distribution is useful, not just descriptive. DRM can use distributional uncertainty to identify unstable reward decisions, and distribution-aware ranking improves Best-of
面壁智能 OpenBMB 发布 Diffusion Reward Models(DRM),不再把人类偏好压缩成单一分数,而是学习完整奖励分布,保留分歧、不确定性与多种可能判断。
来源:X:面壁智能 OpenBMB (@OpenBMB) · x.lingyaoai.com