针对少数样本干扰生成模型对齐,提出自适应修正方法
When Preferences Diverge: Aligning Diffusion Models with Minority-Aware Adaptive DPO
- 引入少数样本感知指标,区分多数与少数偏好数据
- 新损失函数提升主流偏好学习,缓解少数样本负面影响
- 适用于含偏见或不均衡标注的图像生成训练场景
近年来,图像生成领域在微调方法上取得显著进展,尤其在对齐模型与普遍人类偏好方面。本文探讨了偏好数据在扩散模型训练中的关键作用,聚焦于Diffusion-DPO及其后续改进。研究揭示了图像生成中普遍人类偏好的主观性,以及偏好数据集中少数样本带来的挑战。初步实验表明少数样本存在且会损害模型性能。为此,我们提出Adaptive-DPO——一种将少数实例感知度量(包括标注者内部置信度与跨标注者稳定性)融入DPO目标的新方法。该度量可区分多数与少数样本,并设计了自适应损失函数,在两个方面优化DPO:强化模型对多数标签的学习,同时减轻少数样本的负面影响。实验表明,该方法能有效处理合成少数数据与真实世界偏好数据,为图像生成任务提供更高效的训练范式。
原文摘要 · Abstract (English)
In recent years, the field of image generation has witnessed significant advancements, particularly in fine-tuning methods that align models with universal human preferences. This paper explores the critical role of preference data in the training process of diffusion models, particularly in the context of Diffusion-DPO and its subsequent adaptations. We investigate the complexities surrounding universal human preferences in image generation, highlighting the subjective nature of these preferences and the challenges posed by minority samples in preference datasets. Through pilot experiments, we demonstrate the existence of minority samples and their detrimental effects on model performance. We propose Adaptive-DPO -- a novel approach that incorporates a minority-instance-aware metric into the DPO objective. This metric, which includes intra-annotator confidence and inter-annotator stability, distinguishes between majority and minority samples. We introduce an Adaptive-DPO loss function which improves the DPO loss in two ways: enhancing the model's learning of majority labels while mitigating the negative impact of minority samples. Our experiments demonstrate that this method effectively handles both synthetic minority data and real-world preference data, paving the way for more effective training methodologies in image generation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。