arXiv:2512.04332cs.LG2025-12被引 14

用数据正则化解决扩散模型强化学习中的奖励欺骗问题。

Data-regularized Reinforcement Learning for Diffusion Models at Scale

  • 引入前向KL散度将策略锚定在离线数据分布上,提升正则化可靠性。
  • 在百万级GPU小时和万次人工评估中,显著提升人类偏好评分。
  • 适合大规模扩散模型后训练,尤其视频生成场景下表现稳健。

通过强化学习(RL)对齐生成扩散模型与人类偏好至关重要但极具挑战。现有算法常因正则化不可靠而出现奖励欺骗,如质量下降、过度风格化或多样性减少。我们分析表明,这源于其正则化机制的内在缺陷。本文提出数据正则化扩散强化学习(DDRL),利用前向KL散度将策略锚定于离线数据分布。理论上,DDRL实现RL与标准扩散训练的鲁棒、无偏融合;实证上,形成一种简单有效的算法,同时优化奖励与扩散损失。基于超百万GPU小时实验与一万次双盲人工评估,在高分辨率视频生成任务中,DDRL显著提升奖励并缓解基线中的奖励欺骗,取得最高人类偏好评分,建立了一种稳健可扩展的扩散模型后训练范式。

原文摘要 · Abstract (English)

Aligning generative diffusion models with human preferences via reinforcement learning (RL) is critical yet challenging. Most existing algorithms are often vulnerable to reward hacking, such as quality degradation, over-stylization, or reduced diversity. Our analysis demonstrates that this can be attributed to the inherent limitations of their regularization, which provides unreliable penalties. We introduce Data-regularized Diffusion Reinforcement Learning (DDRL), a novel framework that uses the forward KL divergence to anchor the policy to an off-policy data distribution. Theoretically, DDRL enables robust, unbiased integration of RL with standard diffusion training. Empirically, this translates into a simple yet effective algorithm that combines reward maximization with diffusion loss minimization. With over a million GPU hours of experiments and ten thousand double-blind human evaluations, we demonstrate on high-resolution video generation tasks that DDRL significantly improves rewards while alleviating the reward hacking seen in baselines, achieving the highest human preference and establishing a robust and scalable paradigm for diffusion post-training.

扩散模型强化学习视频生成数据正则

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。