arXiv:2605.10937cs.CV2026-05被引 1

通过非线性优势重塑,提升文本到图像模型的后训练效果。

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping

论文配图:Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping
图 1 · 摘自论文原文
  • 从信息几何出发,设计非线性优势重塑机制
  • 在多个基准上超越基线,加速训练并减少奖励欺骗
  • 适合追求生成质量与鲁棒性的模型优化研究者

近期基于强化学习的后训练方法,尤其是组相对策略优化(GRPO),已成为推动文本到图像(T2I)模型发展的主流范式。然而,这些方法易受奖励欺骗影响,模型会利用不完美奖励函数中的偏差而非真正提升性能。本文指出,归一化可能导致校准失真,直接移除提示级标准差项虽能获得线性优势方向的最优策略上升路径,但仍难以分离真实信号与噪声。为此,我们提出超线性优势重塑(SLAS),从信息几何视角重新审视函数更新。通过引入依赖优势的加权费希尔-罗伊信息度量,SLAS构建了非线性几何结构,重塑局部策略空间:在高优势方向放宽约束以放大有效更新,在低优势区域收紧以抑制虚假梯度。同时,采用批级归一化以应对奖励尺度变化。大量实验表明,SLAS在多个骨干模型和基准上持续优于DanceGRPO基线,实现更快训练动态、在GenEval和UniGenBench++上的更优跨域表现、更强的模型缩放鲁棒性,并有效缓解奖励欺骗,保持生成结果的语义与组合保真度。

原文摘要 · Abstract (English)

Recently, post-training methods based on reinforcement learning, with a particular focus on Group Relative Policy Optimization (GRPO), have emerged as the robust paradigm for further advancement of text-to-image (T2I) models. However, these methods are often prone to reward hacking, wherein models exploit biases in imperfect reward functions rather than yielding genuine performance gains. In this work, we identify that normalization could lead to miscalibration and directly removing the prompt-level standard deviation term yields an optimal policy ascent direction that is linear in the advantage but still limits the separation of genuine signals from noise. To mitigate the above issues, we propose Super-Linear Advantage Shaping (SLAS) by revisiting the functional update from an information geometry perspective. By extending the Fisher-Rao information metric with advantage-dependent weighting, SLAS introduces a non-linear geometric structure that reshapes the local policy space. This design relaxes constraints along high-advantage directions to amplify informative updates, while tightening those in low-advantage regions to suppress illusory gradients. In addition, batch-level normalization is applied to stabilize training under varying reward scales. Extensive evaluations demonstrate that SLAS consistently surpasses the DanceGRPO baseline across multiple backbones and benchmarks. In particular, it yields faster training dynamics, improved out-of-domain performance on GenEval and UniGenBench++, and enhanced robustness to model scaling, while mitigating reward hacking and preserving semantic and compositional fidelity in generations.

文本生成强化学习图像生成后训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。