arXiv:2505.18736cs.CV2025-05AAAI被引 2

改进扩散模型的偏好优化,提升生成图像与人类偏好的一致性。

Rethinking Direct Preference Optimization in Diffusion Models

  • 引入可稳定更新的参考模型,平衡探索与优化稳定性。
  • 设计时间步感知训练策略,缓解不同时刻奖励尺度失衡问题。
  • 适用于多种偏好优化算法,适合图像生成与对齐研究者使用。

将文本到图像(T2I)扩散模型与人类偏好对齐已成为关键研究挑战。尽管近期工作已将大语言模型中的偏好优化技术拓展至扩散模型,但普遍存在探索能力有限的问题。本文提出一种新颖且正交的扩散模型偏好优化方法:首先,引入稳定的参考模型更新策略,放宽参考模型冻结限制,通过正则化保持优化锚点稳定,促进探索;其次,提出时间步感知训练策略,缓解不同时间步间奖励尺度失衡问题。该方法可无缝集成于多种偏好优化算法中。实验表明,其显著提升了现有先进方法在人类偏好评估基准上的表现。代码已开源:https://github.com/kaist-cvml/RethinkingDPO_Diffusion_Models。

原文摘要 · Abstract (English)

Aligning text-to-image (T2I) diffusion models with human preferences has emerged as a critical research challenge. While recent advances in this area have extended preference optimization techniques from large language models (LLMs) to the diffusion setting, they often struggle with limited exploration. In this work, we propose a novel and orthogonal approach to enhancing diffusion-based preference optimization. First, we introduce a stable reference model update strategy that relaxes the frozen reference model, encouraging exploration while maintaining a stable optimization anchor through reference model regularization. Second, we present a timestep-aware training strategy that mitigates the reward scale imbalance problem across timesteps. Our method can be integrated into various preference optimization algorithms. Experimental results show that our approach improves the performance of state-of-the-art methods on human preference evaluation benchmarks. The code is available at the Github: https://github.com/kaist-cvml/RethinkingDPO_Diffusion_Models.

扩散模型偏好优化图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。