解决自回归图像生成中强化学习训练不稳的问题
STAGE: Stable and Generalizable GRPO for Autoregressive Image Generation
- 通过优势/KL重加权缓解无关词元的冲突更新
- 引入熵奖励稳定策略更新,减少对预训练分布的破坏
- 在多个基准上提升图像质量与跨任务泛化能力
强化学习近期被用于改进文本到图像生成,但将现有GRPO算法应用于自回归(AR)图像模型仍具挑战。训练过程不稳定易破坏预训练模型能力,导致收益微弱、图像质量下降及泛化性能差。本文重新审视AR图像生成中的GRPO,发现两大关键问题:无用词元带来的矛盾梯度与不稳定的策略熵动态。为此提出STAGE框架,采用两项针对性解决方案:1)优势/KL重加权,基于相似性缓解冲突更新;2)熵奖励,基于参考模型的熵奖励以稳定学习。该方法减轻词元间冲突并稳定训练,减少对预训练分布的破坏,缓解奖励劫持,从而提升泛化能力并更好迁移到其他基准。多基准实验表明,STAGE在视觉质量、稳定性与跨任务泛化方面均优于基线GRPO。
原文摘要 · Abstract (English)
Reinforcement learning has recently been explored to improve text-to-image generation, yet applying existing GRPO algorithms to autoregressive (AR) image models remains challenging. The instability of the training process easily disrupts the pretrained model capability during long runs, resulting in marginal gains, degraded image quality, and poor generalization. In this work, we revisit GRPO for AR image generation and identify two key issues: contradictory gradients from unnecessary tokens and unstable policy entropy dynamics. To address these, we introduce STAGE, a stable and generalizable framework that leverages two targeted solutions: 1) Advantage/KL reweighting. Similarity-aware reweighting to alleviate conflicting updates; and 2) Entropy reward. An entropy-based reward corresponding to reference model to stabilize learning. With the help of alleviating conflicts between tokens and an entropy reward for stabilizing training, we reduce disruption of the pretrained distribution and mitigate reward hacking, which in turn improves generalization and transfer better to other benchmarks. Experiments across multiple benchmarks show that STAGE consistently improves visual quality, stability, and cross-task generalization compared to baseline GRPO.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。