arXiv:2505.23316cs.CL2025-05NeurIPS被引 7

解决大模型对齐中偏好对比导致的生成概率下降问题

Proximalized Preference Optimization for Diverse Feedback Types: A Decomposed Perspective on DPO

  • 将DPO损失重新分解,揭示其隐含正则项缺失
  • 提出PRO方法,统一处理多种反馈类型且避免概率压缩
  • 在配对、二元和标量反馈上均优于现有方法

直接对齐方法通常通过对比偏好与不偏好响应的似然来训练大语言模型。尽管能捕捉相对偏好,但普遍观察到会抑制具体响应的绝对似然,导致模型偏离预期模式,甚至出现奖励劫持现象。我们称此为似然不足确定性,由此重新审视经典直接偏好优化(DPO)。研究发现DPO损失可被合理分解,重构后的损失不仅自然扩展至多种反馈类型,还揭示了似然不足的根本原因:标准DPO隐式简化了重构损失中的正则项。恢复完整正则项可有效解决该问题。基于此,我们提出近端偏好优化(PRO),一种统一的对齐方法,在高效近似完整正则项的基础上,兼容多样反馈类型并消除似然不足。实验表明,PRO在配对、二元及标量反馈场景下均持续优于现有方法。

原文摘要 · Abstract (English)

Direct alignment methods typically train large language models (LLMs) by contrasting the likelihoods of preferred and dispreferred responses. While effective at capturing relative preferences, these methods are widely observed to suppress the absolute likelihoods of example responses. As a result, aligned models can deviate from expected patterns, exhibiting rewar-hacking effect even without an explicit reward model. This fundamental limitation of contrastive alignment, which we term likelihood underdetermination, motivates us to revisit direct preference optimization (DPO) -- the seminal direct alignment method. Interestingly, we show that the DPO loss admits a principled decomposition. The reformulated loss not only extends naturally to a broader range of feedback types, but also unveils the root cause of likelihood underdetermination. Specifically, we identify that standard DPO implicitly oversimplifies a regularizer in the reformulated loss; restoring this full term effectively resolves the underdetermination. Building on these insights, we introduce PRoximalized PReference Optimization (PRO), a unified alignment method that accommodates diverse feedback types while eliminating likelihood underdetermination through an efficient approximation of the full regularizer. Empirical evaluations demonstrate the consistent superiority of PRO over existing methods across pairwise, binary and scalar feedback.

大模型对齐偏好优化生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。