arXiv:2505.11245cs.CV2025-05ICLR被引 26

让扩散模型学会避开不良输出,提升生成结果的偏好一致性。

Diffusion-NPO: Negative Preference Optimization for Better Preference Aligned Generation of Diffusion Models

  • 专门训练模型识别负向偏好,优化生成质量。
  • 无需新数据或训练策略,适配多种主流模型。
  • 显著提升图像与视频生成的用户偏好对齐度。

扩散模型在图像生成方面取得了显著进展,但基于大规模未过滤数据集训练的模型常生成与人类偏好不符的内容。尽管已有诸多方法用于微调预训练扩散模型以提升偏好对齐性,我们指出现有方法忽略了无条件/负条件输出处理的关键作用,导致分类器无关引导(CFG)的有效性下降。为此,我们提出一种简单而通用的方法:专门训练模型感知负向偏好。该方法无需新增训练策略或数据集,仅需对现有技术进行小幅调整。我们的方法可无缝集成至SD1.5、SDXL、视频扩散模型及已进行偏好优化的模型中,持续提升其与人类偏好的对齐程度。

原文摘要 · Abstract (English)

Diffusion models have made substantial advances in image generation, yet models trained on large, unfiltered datasets often yield outputs misaligned with human preferences. Numerous methods have been proposed to fine-tune pre-trained diffusion models, achieving notable improvements in aligning generated outputs with human preferences. However, we argue that existing preference alignment methods neglect the critical role of handling unconditional/negative-conditional outputs, leading to a diminished capacity to avoid generating undesirable outcomes. This oversight limits the efficacy of classifier-free guidance~(CFG), which relies on the contrast between conditional generation and unconditional/negative-conditional generation to optimize output quality. In response, we propose a straightforward but versatile effective approach that involves training a model specifically attuned to negative preferences. This method does not require new training strategies or datasets but rather involves minor modifications to existing techniques. Our approach integrates seamlessly with models such as SD1.5, SDXL, video diffusion models and models that have undergone preference optimization, consistently enhancing their alignment with human preferences.

扩散模型偏好对齐负向优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。