arXiv:2602.09533cs.AI2026-02被引 1

改进大模型对齐方法,让生成更符合人类偏好

Autoregressive Direct Preference Optimization

  • 从生成过程出发重构偏好优化框架,显式引入自回归假设
  • 推导出新损失函数,使计算更简洁且理论更严密
  • 首次区分两种长度度量,为算法设计提供新思路

直接偏好优化(DPO)是将大语言模型与人类偏好对齐的有前景方法。然而,其广泛依赖的响应级Bradley-Terry(BT)模型可能限制了其潜力,因为参考模型和可学习模型仅在目标函数推导后才被假定为自回归。受此限制启发,我们重新审视DPO的理论基础,提出一种新公式,在应用BT模型前显式引入自回归假设。通过重构和扩展DPO,我们推导出一种新变体——自回归直接偏好优化(ADPO),该方法将自回归建模明确整合进偏好优化框架。在不违背理论基础的前提下,推导出的损失函数形式优雅:将DPO目标中求和操作移出log-sigmoid函数。此外,通过对ADPO的理论分析,我们发现设计基于DPO的算法时需考虑两个长度度量:标记长度μ和反馈长度μ'。据我们所知,这是首个显式区分并分析这两个度量对大语言模型偏好优化影响的工作。

原文摘要 · Abstract (English)

Direct preference optimization (DPO) has emerged as a promising approach for aligning large language models (LLMs) with human preferences. However, the widespread reliance on the response-level Bradley-Terry (BT) model may limit its full potential, as the reference and learnable models are assumed to be autoregressive only after deriving the objective function. Motivated by this limitation, we revisit the theoretical foundations of DPO and propose a novel formulation that explicitly introduces the autoregressive assumption prior to applying the BT model. By reformulating and extending DPO, we derive a novel variant, termed Autoregressive DPO (ADPO), that explicitly integrates autoregressive modeling into the preference optimization framework. Without violating the theoretical foundations, the derived loss takes an elegant form: it shifts the summation operation in the DPO objective outside the log-sigmoid function. Furthermore, through theoretical analysis of ADPO, we show that there exist two length measures to be considered when designing DPO-based algorithms: the token length $μ$ and the feedback length $μ'$. To the best of our knowledge, we are the first to explicitly distinguish these two measures and analyze their implications for preference optimization in LLMs.

大模型对齐偏好优化自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。