arXiv:2505.21893cs.LGcs.AI2025-05

提出SIPO框架,解决扩散模型对齐中的训练不稳与策略偏差问题。

SIPO: Stabilized and Improved Preference Optimization for Aligning Diffusion Models

  • 通过梯度裁剪与无效时间步屏蔽稳定训练过程。
  • 引入时间步感知重加权,缓解策略分布偏移问题。
  • 在多个图像和视频生成模型上表现更优,适合需要稳定对齐的场景。

偏好学习已成为对齐扩散模型与人类视觉生成偏好的有效技术。然而,现有方法如Diffusion-DPO面临两大挑战:不同时间步高梯度方差导致的训练不稳定性及参数敏感性,以及优化数据与策略模型分布差异引发的离策略偏差。本文系统分析了不同时间步的扩散轨迹,发现不稳定性主要源于重要性权重较低的早期时间步。为此,提出一种稳定且改进的偏好优化框架SIPO。核心是引入DPO-C&M梯度,通过裁剪与屏蔽无信息时间步以稳定训练;随后采用时间步感知的重要性重加权机制,缓解离策略偏差并强化整个对齐过程中的有效更新。在SD1.5、SDXL、CogVideoX-2B/5B、Wan2.1-1.3B等多类基线模型上的广泛实验表明,SIPO能持续提升训练稳定性,优于需精细调参的现有对齐方法。结果表明时间步感知对齐至关重要,为改进扩散模型偏好优化提供了实用指导。

原文摘要 · Abstract (English)

Preference learning has garnered extensive attention as an effective technique for aligning diffusion models with human preferences in visual generation. However, existing alignment approaches such as Diffusion-DPO suffer from two fundamental challenges: training instability caused by high gradient variances at various timesteps and high parameter sensitivities, and off-policy bias arising from the discrepancy between the optimization data and the policy models' distribution. Our first contribution is a systematic analysis of diffusion trajectories across different timesteps, identifying that the instability primarily originates from early timesteps with low importance weights. To address these issues, we propose \textbf{SIPO}, a \textbf{S}tabilized and \textbf{I}mproved \textbf{P}reference \textbf{O}ptimization framework for aligning diffusion models with human preferences. Concretely, a key gradient, \emph{i.e.,} DPO-C\&M is introduced to stabilize training by clipping and masking uninformative timesteps. This is followed by a timestep-aware importance-reweighting paradigm to mitigate off-policy bias and emphasize informative updates throughout the alignment process. Extensive experiments on various baseline models including image generation models on SD1.5, SDXL, and video generation models CogVideoX-2B/5B, Wan2.1-1.3B, demonstrate that our SIPO consistently promotes stabilized training and outperforms existing alignment methods that with meticulous adjustments on parameters.Overall, these results suggest the importance of timestep-aware alignment and provide valuable guidelines for improved preference optimization in aligning diffusion models.

扩散模型偏好优化训练稳定对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。