arXiv:2505.12435cs.LGcs.AI2025-05ACL被引 5

提出SGDPO算法,提升语言模型对齐人类偏好的能力。

SGDPO: Self-Guided Direct Preference Optimization for Language Model Alignment

  • 用引导项控制梯度流,精细调节优选与拒选响应的更新。
  • 在多个基准上实现最高9.19%的性能提升。
  • 适合关注大模型对齐优化的研究者与实践者。

直接偏好优化(DPO)因灵活性被广泛用于对齐大型语言模型与人类价值观。尽管有效,但其生成人类偏好评价响应的能力有限,结果也缺乏鲁棒性。为此,本文提出一种新型自引导直接偏好优化算法——SGDPO,引入先导项以引导优化过程中的梯度流动,实现对优选与拒选奖励更新的细粒度控制。我们提供了该方法的详细理论分析,并阐明其运作机制。此外,在多种模型和基准上进行了全面实验。大量实验证明,实证结果与理论分析一致,验证了所提方法的有效性,性能最高提升达9.19%。

原文摘要 · Abstract (English)

Direct Preference Optimization (DPO) is broadly utilized for aligning Large Language Models (LLMs) with human values because of its flexibility. Despite its effectiveness, it has been observed that the capability of DPO to generate human-preferred response is limited and the results of DPO are far from resilient. To address these limitations, in this paper we propose a novel Self-Guided Direct Preference Optimization algorithm, i.e., SGDPO, which incorporates a pilot term to steer the gradient flow during the optimization process, allowing for fine-grained control over the updates of chosen and rejected rewards. We provide a detailed theoretical analysis of our proposed method and elucidate its operational mechanism. Furthermore, we conduct comprehensive experiments on various models and benchmarks. The extensive experimental results demonstrate the consistency between the empirical results and our theoretical analysis and confirm the effectiveness of our proposed approach (up to 9.19% higher score).

大模型对齐偏好优化SGDPO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。