arXiv:2506.12725cs.AIcs.CL2025-06EMNLP被引 13

改进DPO损失函数,让模型更专注生成优选回复。

Rethinking DPO: The Role of Rejected Responses in Preference Misalignment

  • 用约束机制限制差回复对损失的影响
  • 在多个数据集上优于现有方法,平衡优化优劣回应
  • 适合需要精准偏好控制的对话系统研究者

直接偏好优化(DPO)是一种简单高效的框架,受到广泛关注。然而,由于被拒回复在损失函数中占主导地位,DPO常难以实现其核心目标——提高优选回复的生成概率,同时降低被拒回复的概率。这种不平衡导致对优选回复的促进效果不佳。本文系统分析了DPO及其相关算法的局限性,并提出一种新方法Bounded-DPO(BDPO),通过约束被拒回复的影响,在保持DPO原始优化结构的同时实现对优选与被拒回复的均衡优化。理论分析和实证结果表明,BDPO在多个基准数据集上均优于现有算法,显著提升性能。

原文摘要 · Abstract (English)

Direct Preference Optimization (DPO) is a simple and efficient framework that has attracted substantial attention. However, it often struggles to meet its primary objectives -- increasing the generation probability of chosen responses while reducing that of rejected responses -- due to the dominant influence of rejected responses on the loss function. This imbalance leads to suboptimal performance in promoting preferred responses. In this work, we systematically analyze the limitations of DPO and existing algorithms designed to achieve the objectives stated above. To address these limitations, we propose Bounded-DPO (BDPO), a novel method that bounds the influence of rejected responses while maintaining the original optimization structure of DPO. Through theoretical analysis and empirical evaluations, we demonstrate that BDPO achieves a balanced optimization of the chosen and rejected responses, outperforming existing algorithms.

偏好优化DPO对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。