arXiv:2506.02698cs.CV2025-06ICML被引 15

用平滑偏好分布提升扩散模型对多样人类偏好的对齐效果。

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences

  • 用奖励模型生成平滑偏好分布,替代二值偏好,增强对个体差异的建模。
  • 通过反向重建技术估算扩散路径偏好分布,实现更精准的优化目标对齐。
  • 在更低训练成本下超越基线,适合需个性化生成的场景。

直接偏好优化(DPO)利用成对偏好数据将文本到图像(T2I)生成模型与人类偏好对齐。尽管数据收集与标注耗费大量资源,但一个关键问题常被忽视:人类偏好存在个体差异,应以更细粒度的方式表示。为此,我们提出SmPO-Diffusion,一种建模偏好分布的新方法,以改进DPO目标,并提供扩散优化目标的数值上界估计。首先,我们引入平滑偏好分布替代原始二值分布,利用奖励模型模拟人类偏好,通过偏好似然平均化改进DPO损失,使当偏好相近时损失趋近于零。此外,我们采用反向重建技术模拟扩散模型的轨迹偏好分布,实现更精确的对齐。该方法通过简单修改有效缓解了现有方法中过度优化和目标错位的问题。SmPO-Diffusion在偏好评估中达到当前最优性能,在多个指标上优于基线,且训练成本更低。

原文摘要 · Abstract (English)

Direct Preference Optimization (DPO) aligns text-to-image (T2I) generation models with human preferences using pairwise preference data. Although substantial resources are expended in collecting and labeling datasets, a critical aspect is often neglected: \textit{preferences vary across individuals and should be represented with more granularity.} To address this, we propose SmPO-Diffusion, a novel method for modeling preference distributions to improve the DPO objective, along with a numerical upper bound estimation for the diffusion optimization objective. First, we introduce a smoothed preference distribution to replace the original binary distribution. We employ a reward model to simulate human preferences and apply preference likelihood averaging to improve the DPO loss, such that the loss function approaches zero when preferences are similar. Furthermore, we utilize an inversion technique to simulate the trajectory preference distribution of the diffusion model, enabling more accurate alignment with the optimization objective. Our approach effectively mitigates issues of excessive optimization and objective misalignment present in existing methods through straightforward modifications. Our SmPO-Diffusion achieves state-of-the-art performance in preference evaluation, outperforming baselines across metrics with lower training costs. The project page is https://jaydenlyh.github.io/SmPO-project-page/.

扩散模型偏好对齐个性化生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。