arXiv:2508.18312cs.LGcs.AI2025-08NeurIPS被引 11

选中回复质量比拒绝回复更重要,直接影响模型对齐效果。

What Matters in Data for DPO?

  • 优先提升选中样本质量,可显著优化对齐效果。
  • 实验验证:选中样本越优,模型性能越强,与拒接样本无关。
  • 适合构建高效对齐数据集的研究者和工程师参考。

直接偏好优化(DPO)已成为一种无需学习奖励模型即可对齐大语言模型与人类偏好的简单有效方法。尽管应用广泛,但关于偏好数据的关键特性仍不明确。本文从理论与实证双重视角系统研究了偏好数据分布对DPO的影响。结果表明,选中回复的质量在优化目标中起主导作用,而拒绝回复的质量影响较小。理论分析揭示了最优响应分布,并说明对比性主要通过提升选中样本质量实现。在线DPO设置下,其等价于仅对选中样本进行监督微调。跨任务的大量实验验证:无论拒绝样本质量如何,提升选中样本质量始终能提高模型表现。此外,我们研究了策略数据混合的收益。结果解释了常见实践机制,为构建高效偏好数据集提供实用指导。

原文摘要 · Abstract (English)

Direct Preference Optimization (DPO) has emerged as a simple and effective approach for aligning large language models (LLMs) with human preferences, bypassing the need for a learned reward model. Despite its growing adoption, a fundamental question remains open: what characteristics of preference data are most critical for DPO performance? In this work, we provide a systematic study of how preference data distribution influences DPO, from both theoretical and empirical perspectives. We show that the quality of chosen responses plays a dominant role in optimizing the DPO objective, while the quality of rejected responses may have relatively limited impact. Our theoretical analysis characterizes the optimal response distribution under DPO and reveals how contrastiveness between responses helps primarily by improving the chosen samples. We further study an online DPO setting and show it effectively reduces to supervised fine-tuning on the chosen responses. Extensive experiments across diverse tasks confirm our findings: improving the quality of chosen responses consistently boosts performance regardless of the quality of the rejected responses. We also investigate the benefit of mixing the on-policy data. Our results interpret the mechanism behind some widely adopted strategies and offer practical insights for constructing high-impact preference datasets for LLM alignment.

大模型对齐偏好优化数据质量训练策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。