arXiv:2502.20847cs.LG2025-02被引 6

DPO训练时梯度不平衡导致性能下降,新方法有效缓解此问题。

Gradient Imbalance in Direct Preference Optimization

  • 通过梯度重加权修正DPO的优化方向
  • 实验验证新方法显著提升模型收敛效果
  • 适合关注强化学习优化稳定性的研究者

直接偏好优化(DPO)被视作基于人类反馈的强化学习(RLHF)中替代近端策略优化(PPO)的有前景方案。然而,实证评估始终显示DPO性能低于主流RLHF流程。本文系统分析DPO训练动态,发现梯度不平衡是关键限制因素。理论与实证均表明,该不平衡会扰动优化轨迹、加剧学习不稳定性并导致次优收敛。为此,我们提出Balanced-DPO,一种简单有效的改进方法,引入计算高效的梯度重加权机制。实验验证了该方法的有效性,证实缓解梯度不平衡对提升DPO性能至关重要,指明了未来研究的重要方向。

原文摘要 · Abstract (English)

Direct Preference Optimization (DPO) has been proposed as a promising alternative to Proximal Policy Optimization (PPO) based Reinforcement Learning with Human Feedback (RLHF). However, empirical evaluations consistently reveal suboptimal performance in DPO compared to common RLHF pipelines. In this work, we conduct a systematic analysis of DPO's training dynamics and identify gradient imbalance as a critical limitation. We demonstrate theoretically and empirically that this imbalance perturbs optimization trajectories, destabilizes learning, and induces suboptimal convergence. To address this issue, we propose Balanced-DPO, a simple yet effective modification to the DPO objective that introduces a computationally efficient gradient reweighting mechanism. Our experiments demonstrate the effectiveness of Balanced-DPO, validating the theoretical findings and confirming that addressing gradient imbalance is key to improving DPO's performance, highlighting a promising direction for future research.

强化学习偏好优化梯度平衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。