arXiv:2511.19049cs.CV2025-11被引 2

解决扩散模型视频生成中的偏好对齐失效问题

Beyond Reward Margin: Rethinking and Resolving Likelihood Displacement in Diffusion Models via Video Generation

  • 提出PG-DPO方法,结合自适应拒收缩放与隐式偏好正则化
  • 实验证明在视频生成任务中优于现有方法,提升生成质量
  • 适合关注扩散模型偏好对齐与视频生成的研究者

直接偏好优化(DPO)在对齐生成结果与人类偏好方面表现良好,但存在似然位移问题:选定样本的概率在训练中反而下降,影响生成质量。尽管该问题在自回归模型中已有研究,但在扩散模型中的影响仍不明确,尤其在视频生成任务中表现不佳。本文在扩散框架下对DPO损失进行形式化分析,揭示了两种失败模式:(1) 优化冲突,源于选定与拒绝样本间奖励差距过小;(2) 次优最大化,由奖励差距过大引起。基于此,提出新型方法PG-DPO,融合自适应拒收缩放(ARS)与隐式偏好正则化(IPR),有效缓解似然位移。实验表明,PG-DPO在定量指标与定性评估中均优于现有方法,为视频生成中的偏好对齐提供了稳健解决方案。

原文摘要 · Abstract (English)

Direct Preference Optimization (DPO) has shown promising results in aligning generative outputs with human preferences by distinguishing between chosen and rejected samples. However, a critical limitation of DPO is likelihood displacement, where the probabilities of chosen samples paradoxically decrease during training, undermining the quality of generation. Although this issue has been investigated in autoregressive models, its impact within diffusion-based models remains largely unexplored. This gap leads to suboptimal performance in tasks involving video generation. To address this, we conduct a formal analysis of DPO loss through updating policy within the diffusion framework, which describes how the updating of specific training samples influences the model's predictions on other samples. Using this tool, we identify two main failure modes: (1) Optimization Conflict, which arises from small reward margins between chosen and rejected samples, and (2) Suboptimal Maximization, caused by large reward margins. Informed by these insights, we introduce a novel solution named Policy-Guided DPO (PG-DPO), combining Adaptive Rejection Scaling (ARS) and Implicit Preference Regularization (IPR) to effectively mitigate likelihood displacement. Experiments show that PG-DPO outperforms existing methods in both quantitative metrics and qualitative evaluations, offering a robust solution for improving preference alignment in video generation tasks.

扩散模型视频生成偏好对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。