通过双标签优化提升扩散模型图像生成质量与提示相关性。
Dual Caption Preference Optimization for Diffusion Models
- 为每对图像分配不同描述,增强偏好学习信号。
- 在多项指标上超越SD 2.1、Diffusion-DPO等基线模型。
- 适合关注生成质量与提示对齐的研究者与开发者。
近期人类偏好优化技术在大型语言模型中的成功,展现出改善文本到图像扩散模型的巨大潜力。这类方法旨在学习优选样本的分布并区分其与次优样本。然而,现有偏好数据集中原始描述往往无法清晰体现优选图像的优势,削弱了训练时的监督信号。为此,我们提出双标签偏好优化(DCPO),一种数据增强与优化框架,通过为每对偏好样本分配两个不同描述,强化模型在训练中区分优选与次优结果的能力。我们构建了改进版的Pick-Double Caption数据集,采用独立描述每张图像,并提出三种生成差异描述的方法:直接标注、扰动和混合策略。实验表明,DCPO在多个指标(包括Pickscore、HPSv2.1、GenEval、CLIPscore、ImageReward)上显著优于基于SD 2.1微调的Stable Diffusion 2.1、SFT_Chosen、Diffusion-DPO和MaPO。
原文摘要 · Abstract (English)
Recent advancements in human preference optimization, originally developed for Large Language Models (LLMs), have shown significant potential in improving text-to-image diffusion models. These methods aim to learn the distribution of preferred samples while distinguishing them from less preferred ones. However, within the existing preference datasets, the original caption often does not clearly favor the preferred image over the alternative, which weakens the supervision signal available during training. To address this issue, we introduce Dual Caption Preference Optimization (DCPO), a data augmentation and optimization framework that reinforces the learning signal by assigning two distinct captions to each preference pair. This encourages the model to better differentiate between preferred and less-preferred outcomes during training. We also construct Pick-Double Caption, a modified version of Pick-a-Pic v2 with separate captions for each image, and propose three different strategies for generating distinct captions: captioning, perturbation, and hybrid methods. Our experiments show that DCPO significantly improves image quality and relevance to prompts, outperforming Stable Diffusion (SD) 2.1, SFT_Chosen, Diffusion-DPO, and MaPO across multiple metrics, including Pickscore, HPSv2.1, GenEval, CLIPscore, and ImageReward, fine-tuned on SD 2.1 as the backbone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。