跳到正文
原文
X:Arena (@arena)· @arena·· 18 小时前AI 评分57

Arena 研究:前沿图像模型后训练需复合奖励,FLUX.2-dev 提升 69 Elo

How to design rewards for post-training frontier image models? Our research suggests human preference reward is necessary, but insufficient: A preference model may still reward outputs that look appealing but miss details, introduce unrequested content, or exhibit other forms of reward-hacking. We therefore optimize towards a composite reward: - Bradley-Terry reward model trained on ~5.6M pairwise human votes - Faithfulness reward from auto-generated prompt checklists evaluated by a vision-language model - Constraint reward covering explicit and implicit user intent - Anti-reward-hacking rubric rewards targeting failures such as garbled text and photorealism drift This post-training recipe improves two already-strong open image models: - Post-trained FLUX.2-dev gains 69 Elo points on our live T2I leaderboard, scoring 1202 - Post-trained Ideogram 4 gains 20 Elo points reaching a score of 1224 and surpassing all publicly listed open models (as of Sep 04, 2026). Offline ablations with Gem

AI 导读

Arena 研究提出前沿图像模型后训练的复合奖励方案,认为人类偏好奖励必要但不充分,偏好模型仍可能奖励细节缺失、引入未请求内容等 reward-hacking 行为。

来源:X:Arena (@arena) · x.lingyaoai.com