arXiv:2508.20751cs.CV2025-08被引 96

用偏好比较替代分数评分,让文生图强化学习更稳定。

Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning

  • 用成对图像偏好代替单图评分,避免奖励作弊
  • 在600个提示下验证,显著提升生成稳定性
  • 适合关注文生图训练稳定性的研究者

近期研究表明,基于GRPO的强化学习方法在提升文生图(T2I)生成质量方面具有重要意义。然而,当前采用点式奖励模型(RM)对生成图像打分的方法易受奖励黑客攻击。我们发现,当图像间评分差异极小时,归一化会放大这些微小差距,导致模型过度优化于无关紧要的细微差别,最终破坏生成过程的稳定性。为此,我们提出Pref-GRPO,一种基于成对偏好奖励的GRPO方法,将优化目标从分数最大化转变为偏好拟合,实现更稳定的训练。在该方法中,每组图像通过偏好RM进行两两比较,以胜率作为奖励信号。大量实验表明,Pref-GRPO能有效区分细微的图像质量差异,提供更可靠的稳定优势,缓解奖励黑客问题。此外,现有T2I评估基准因评价标准粗糙而难以全面评估模型性能。为此,我们构建了UniGenBench——一个包含600个提示、5大主题与20个子主题的统一评测基准。该基准通过10项主准则与27项子准则评估语义一致性,利用多模态大模型(MLLM)完成基准构建与评估。实验证明,该基准可揭示开源与闭源T2I模型的优劣势,并验证Pref-GRPO的有效性。

原文摘要 · Abstract (English)

Recent advancements highlight the importance of GRPO-based reinforcement learning methods and benchmarking in enhancing text-to-image (T2I) generation. However, current methods using pointwise reward models (RM) for scoring generated images are susceptible to reward hacking. We reveal that this happens when minimal score differences between images are amplified after normalization, creating illusory advantages that drive the model to over-optimize for trivial gains, ultimately destabilizing the image generation process. To address this, we propose Pref-GRPO, a pairwise preference reward-based GRPO method that shifts the optimization objective from score maximization to preference fitting, ensuring more stable training. In Pref-GRPO, images are pairwise compared within each group using preference RM, and the win rate is used as the reward signal. Extensive experiments demonstrate that PREF-GRPO differentiates subtle image quality differences, providing more stable advantages and mitigating reward hacking. Additionally, existing T2I benchmarks are limited by coarse evaluation criteria, hindering comprehensive model assessment. To solve this, we introduce UniGenBench, a unified T2I benchmark comprising 600 prompts across 5 main themes and 20 subthemes. It evaluates semantic consistency through 10 primary and 27 sub-criteria, leveraging MLLM for benchmark construction and evaluation. Our benchmarks uncover the strengths and weaknesses of both open and closed-source T2I models and validate the effectiveness of Pref-GRPO.

文生图强化学习奖励设计稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。