arXiv:2605.21123cs.CVcs.LG2026-05被引 2

改进扩散与流匹配模型的对齐方法,提升文本生成图像质量。

Linear-DPO: Linear Direct Preference Optimization for Diffusion and Flow-Matching Generative Models

  • 基于统一反向SDE框架,推导出覆盖两类生成模型的通用对齐目标。
  • 在SD1.5、SDXL和SD3-Medium上均优于现有基线,生成质量显著提升。
  • 适合从事文本到图像生成、模型对齐研究的科研人员参考。

直接偏好优化(DPO)在大语言模型对齐中表现优异,但在文本到图像生成任务中仍面临挑战。现有研究仅局限于去噪扩散模型,忽略了流匹配模型,且将基于离散NLP的DPO应用于回归型生成任务时存在目标不匹配问题。本文通过统一的反向时间随机微分方程(SDE)框架,推导出涵盖扩散与流匹配模型的广义DPO目标,并从梯度角度指出标准DPO在文本到图像生成中次优。为此,我们提出Linear-DPO,用持续线性效用函数替代激进的sigmoid效用,结合EMA更新的参考模型。在扩散模型(SD1.5、SDXL)和流匹配模型(SD3-Medium)上的定性和定量实验表明,该方法优于现有基线。

原文摘要 · Abstract (English)

Direct Preference Optimization (DPO) is successful for alignment in LLMs but still faces challenges in text-to-image generation. Existing studies are confined to denoising diffusion models while overlooking flow-matching, and suffer from an objective mismatch when applying discrete NLP-based DPO to regression-based generative tasks.\ In this paper, we derive a generalized DPO objective that covers both diffusion and flow-matching via a unified reverse-time SDE framework, and point out from a gradient perspective that the standard DPO objective is suboptimal for text-to-image generation. Consequently, we propose Linear-DPO, which replaces the aggressive sigmoid-based utility function with a sustained linear utility and incorporates an EMA-updated reference model. Qualitative and quantitative experiments on diffusion models (SD1.5, SDXL) and flow-matching model (SD3-Medium) demonstrate the superiority of our approach over existing baselines.

生成模型对齐优化扩散模型流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。