用合成反馈提升扩散模型对齐,成本低且效果不输人工标注。
DeDPO: Debiased Direct Preference Optimization for Diffusion Models
- 引入因果推断的去偏技术,修正合成反馈中的系统性偏差。
- 在少量人工标注下,性能媲美甚至超越全人工数据训练的模型。
- 适合需要低成本大规模对齐的扩散模型研究与应用者。
直接偏好优化(DPO)已成为扩散模型对齐的主流方法,可在无需显式奖励建模的情况下实现离策略训练。然而,其对大规模高质量人工偏好标签的依赖带来了严重成本与可扩展性瓶颈。为此,我们提出一种半监督框架,通过低成本的合成AI反馈扩充有限的人工数据。本文提出去偏直接偏好优化(DeDPO),首次将因果推断中的去偏估计技术融入DPO目标函数。通过显式识别并纠正合成标注器固有的系统性偏差与噪声,DeDPO确保了从不完美反馈源(如自训练和视觉-语言模型)中稳健学习。实验表明,DeDPO对合成标注方法的变化具有鲁棒性,在性能上达到甚至偶尔超过完全人工标注数据训练模型的理论上限。这确立了DeDPO作为使用廉价合成监督实现人机对齐的可扩展解决方案。
原文摘要 · Abstract (English)
Direct Preference Optimization (DPO) has emerged as a predominant alignment method for diffusion models, facilitating off-policy training without explicit reward modeling. However, its reliance on large-scale, high-quality human preference labels presents a severe cost and scalability bottleneck. To overcome this, We propose a semi-supervised framework augmenting limited human data with a large corpus of unlabeled pairs annotated via cost-effective synthetic AI feedback. Our paper introduces Debiased DPO (DeDPO), which uniquely integrates a debiased estimation technique from causal inference into the DPO objective. By explicitly identifying and correcting the systematic bias and noise inherent in synthetic annotators, DeDPO ensures robust learning from imperfect feedback sources, including self-training and Vision-Language Models (VLMs). Experiments demonstrate that DeDPO is robust to the variations in synthetic labeling methods, achieving performance that matches and occasionally exceeds the theoretical upper bound of models trained on fully human-labeled data. This establishes DeDPO as a scalable solution for human-AI alignment using inexpensive synthetic supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。