arXiv:2502.03095cs.LG2025-02被引 6

揭示DPO与强化学习算法的内在联系,厘清其是否属于强化学习。

Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms

  • 构建统一框架UDRRA,从损失函数构造解析DPO与RLHF的关系。
  • 发现DPO与PPO收敛至相同目标策略分布,证明其本质为强化学习。
  • 分析关键组件影响,为优化训练效率提供理论依据。

随着大语言模型的快速发展,基于人类反馈的强化学习(RLHF)算法被广泛用于提升模型的安全性和对齐人类偏好。这些算法可分为两类:依赖显式奖励函数的基于演员-评论家的近端策略优化(PPO)和基于对齐的直接偏好优化(DPO)。DPO使用由人类偏好数据驱动的分类损失,与PPO存在差异,引发对其是否属于强化学习的争议。为解决这一困惑,本文聚焦三个核心问题:(1)损失函数的构建方式;(2)算法收敛的目标分布;(3)损失函数中关键组件的影响。我们首先建立统一框架UDRRA,连接各类算法;其次在该框架下揭示其目标策略分布;最后研究DPO中关键组件对收敛速度的影响。本工作深化了对DPO、强化学习及其他RLHF算法之间关系的理解,为改进现有算法提供了新视角。

原文摘要 · Abstract (English)

With the rapid development of Large Language Models (LLMs), numerous Reinforcement Learning from Human Feedback (RLHF) algorithms have been introduced to improve model safety and alignment with human preferences. These algorithms can be divided into two main frameworks based on whether they require an explicit reward (or value) function for training: actor-critic-based Proximal Policy Optimization (PPO) and alignment-based Direct Preference Optimization (DPO). The mismatch between DPO and PPO, such as DPO's use of a classification loss driven by human-preferred data, has raised confusion about whether DPO should be classified as a Reinforcement Learning (RL) algorithm. To address these ambiguities, we focus on three key aspects related to DPO, RL, and other RLHF algorithms: (1) the construction of the loss function; (2) the target distribution at which the algorithm converges; (3) the impact of key components within the loss function. Specifically, we first establish a unified framework named UDRRA connecting these algorithms based on the construction of their loss functions. Next, we uncover their target policy distributions within this framework. Finally, we investigate the critical components of DPO to understand their impact on the convergence rate. Our work provides a deeper understanding of the relationship between DPO, RL, and other RLHF algorithms, offering new insights for improving existing algorithms.

强化学习大模型对齐DPO策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。