arXiv:2411.14871cs.CVcs.AI2024-11被引 3

让扩散模型生成更符合人类偏好的图像,关键在优化中间去噪步骤。

Preference Alignment for Diffusion Model via Explicit Denoised Distribution Estimation

  • 提出显式估计去噪分布的方法,连接中间步骤与最终结果。
  • 两种估计策略使模型在去噪中段优化效果显著提升。
  • 适合追求生成质量与人类偏好对齐的研究者或应用开发者。

扩散模型在文本到图像生成中表现卓越,使其偏好对齐愈发重要。但偏好标签通常仅在去噪轨迹末端可用,难以优化中间步骤。本文提出显式去噪分布估计(DDE),将中间步骤与最终去噪分布直接关联,从而实现全程轨迹优化。设计了两种估计策略:分步估计利用条件去噪分布推断模型输出分布;单次估计通过DDIM建模将模型输出转换为终端去噪分布。理论与实证分析表明,DDE自然导出一种新信用分配机制,优先优化去噪过程的中段。大量实验显示,该方法在定量与定性指标上均优于现有方法。

原文摘要 · Abstract (English)

Diffusion models have shown remarkable success in text-to-image generation, making preference alignment for these models increasingly important. The preference labels are typically available only at the terminal of denoising trajectories, which poses challenges in optimizing the intermediate denoising steps. In this paper, we propose to conduct Denoised Distribution Estimation (DDE) that explicitly connects intermediate steps to the terminal denoised distribution. Therefore, preference labels can be used for the entire trajectory optimization. To this end, we design two estimation strategies for our DDE. The first is stepwise estimation, which utilizes the conditional denoised distribution to estimate the model denoised distribution. The second is single-shot estimation, which converts the model output into the terminal denoised distribution via DDIM modeling. Analytically and empirically, we reveal that DDE equipped with two estimation strategies naturally derives a novel credit assignment scheme that prioritizes optimizing the middle part of the denoising trajectory. Extensive experiments demonstrate that our approach achieves superior performance, both quantitatively and qualitatively.

扩散模型偏好对齐去噪分布

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。