arXiv:2505.22002cs.CV2025-05ICML被引 17

解决扩散模型图文对齐中的视觉不一致问题,提升生成质量。

D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples

  • 通过掩码引导的自注意力融合,生成视觉一致且对齐良好的图像。
  • 保留去噪轨迹,支持直接偏好优化训练,显著提升对齐效果。
  • 适用于多种强化学习算法,适合追求高质量图文生成的研究者。

扩散模型的实际应用受限于生成图像与文本提示之间的对齐不足。尽管近期研究引入直接偏好优化(DPO)以改善对齐,但其效果受制于视觉不一致问题:对齐良好与不良图像间存在显著视觉差异,导致模型在微调时难以识别提升对齐的关键因素。为此,本文提出D-Fusion,一种构建可用于DPO训练的视觉一致样本的方法。一方面,通过掩码引导的自注意力融合,生成的图像不仅对齐良好,且与给定的低质量图像在视觉上保持一致;另一方面,D-Fusion能保留生成图像的去噪轨迹,这对DPO训练至关重要。大量实验表明,D-Fusion在应用于不同强化学习算法时,均有效提升了提示-图像对齐能力。

原文摘要 · Abstract (English)

The practical applications of diffusion models have been limited by the misalignment between generated images and corresponding text prompts. Recent studies have introduced direct preference optimization (DPO) to enhance the alignment of these models. However, the effectiveness of DPO is constrained by the issue of visual inconsistency, where the significant visual disparity between well-aligned and poorly-aligned images prevents diffusion models from identifying which factors contribute positively to alignment during fine-tuning. To address this issue, this paper introduces D-Fusion, a method to construct DPO-trainable visually consistent samples. On one hand, by performing mask-guided self-attention fusion, the resulting images are not only well-aligned, but also visually consistent with given poorly-aligned images. On the other hand, D-Fusion can retain the denoising trajectories of the resulting images, which are essential for DPO training. Extensive experiments demonstrate the effectiveness of D-Fusion in improving prompt-image alignment when applied to different reinforcement learning algorithms.

扩散模型图文对齐生成质量偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。