arXiv:2510.01540cs.CV2025-10被引 8

让扩散模型学会按完整排名优化图像生成偏好

Towards Better Optimization For Listwise Preference in Diffusion Models

  • 基于普莱克特-卢克模型扩展DPO,用完整排名代替成对比较
  • 在多类任务中显著提升图像质量与用户偏好契合度
  • 适合需要精细用户偏好建模的生成式AI研究者

从人类反馈中强化学习(RLHF)已被证明能有效对齐文本到图像(T2I)扩散模型与人类偏好。尽管直接偏好优化(DPO)因计算效率高且无需显式奖励建模而被广泛应用,但其在扩散模型中的应用主要依赖成对偏好。对列表级偏好的精确优化仍缺乏系统研究。实践中,人类对图像的反馈常包含隐含的排序信息,比成对比较更能准确反映偏好。本文提出Diffusion-LPO,一种针对扩散模型的简单高效列表级偏好优化框架。给定一段描述,将用户反馈聚合为图像排序列表,并在普莱克特-卢克模型下推导出DPO的目标扩展。Diffusion-LPO通过强制每个样本优于所有排名更低的样本,确保整个排序的一致性。我们在文本到图像生成、图像编辑和个性化偏好对齐等多个任务上实证验证了其有效性。Diffusion-LPO在视觉质量和偏好对齐方面均持续优于成对DPO基线。

原文摘要 · Abstract (English)

Reinforcement learning from human feedback (RLHF) has proven effectiveness for aligning text-to-image (T2I) diffusion models with human preferences. Although Direct Preference Optimization (DPO) is widely adopted for its computational efficiency and avoidance of explicit reward modeling, its applications to diffusion models have primarily relied on pairwise preferences. The precise optimization of listwise preferences remains largely unaddressed. In practice, human feedback on image preferences often contains implicit ranked information, which conveys more precise human preferences than pairwise comparisons. In this work, we propose Diffusion-LPO, a simple and effective framework for Listwise Preference Optimization in diffusion models with listwise data. Given a caption, we aggregate user feedback into a ranked list of images and derive a listwise extension of the DPO objective under the Plackett-Luce model. Diffusion-LPO enforces consistency across the entire ranking by encouraging each sample to be preferred over all of its lower-ranked alternatives. We empirically demonstrate the effectiveness of Diffusion-LPO across various tasks, including text-to-image generation, image editing, and personalized preference alignment. Diffusion-LPO consistently outperforms pairwise DPO baselines on visual quality and preference alignment.

扩散模型偏好优化生成式AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。