arXiv:2502.01667cs.LGcs.AI2025-02被引 8

改进扩散模型对齐方法,让中间生成步骤更符合人类审美偏好。

Refining Alignment Framework for Diffusion Models with Intermediate-Step Preference Ranking

  • 基于每步噪声样本的奖励排序,直接优化中间阶段
  • 实验显示生成图像在美学和偏好度上显著提升
  • 适合关注生成质量与人类偏好对齐的研究者

直接偏好优化(DPO)在对齐扩散模型与人类偏好方面表现优异。以往方法通常假设最终生成结果与中间噪声样本具有一致的偏好标签,并直接将DPO应用于这些噪声样本进行微调。然而,我们从梯度方向和偏好顺序两个角度理论揭示了该假设的内在问题及其对偏好对齐效果的影响。为此,我们提出一种针对性的偏好优化框架(TailorPO),基于逐步奖励对中间噪声样本进行直接排序,并通过简洁高效的设计有效解决梯度方向偏差问题。此外,我们将扩散模型的梯度引导融入偏好对齐过程,进一步提升优化效率。实验结果表明,该方法显著增强了模型生成美观且受人类偏好的图像的能力。

原文摘要 · Abstract (English)

Direct preference optimization (DPO) has shown success in aligning diffusion models with human preference. Previous approaches typically assume a consistent preference label between final generations and noisy samples at intermediate steps, and directly apply DPO to these noisy samples for fine-tuning. However, we theoretically identify inherent issues in this assumption and its impacts on the effectiveness of preference alignment. We first demonstrate the inherent issues from two perspectives: gradient direction and preference order, and then propose a Tailored Preference Optimization (TailorPO) framework for aligning diffusion models with human preference, underpinned by some theoretical insights. Our approach directly ranks intermediate noisy samples based on their step-wise reward, and effectively resolves the gradient direction issues through a simple yet efficient design. Additionally, we incorporate the gradient guidance of diffusion models into preference alignment to further enhance the optimization effectiveness. Experimental results demonstrate that our method significantly improves the model's ability to generate aesthetically pleasing and human-preferred images.

扩散模型偏好对齐图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。