arXiv:2409.03650cs.LGcs.CL2024-09EMNLP被引 29

DPO学习的隐式奖励模型泛化能力弱于显式模型,尤其在分布外数据上表现差。

On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization

  • 比较DPO与显式奖励模型在偏好判断上的表现差异
  • 在5个分布外设置中,DPO准确率平均下降3%,最高降7%
  • 提示迭代DPO应引入显式奖励模型以提升鲁棒性

基于人类反馈的强化学习(RLHF)是使语言模型对齐人类偏好的有效方法,核心在于学习一个用于评分人类偏好的奖励函数。主流方法包括训练显式奖励模型(EXRM)和通过直接偏好优化(DPO)从偏好数据中学习隐式奖励。尽管理论表明DPO学习的隐式奖励模型(DPORM)在极限下可逼近EXRM,且其有效性暗示了最优策略学习,但其在实际中的表现仍不明确。本文研究了DPORM与EXRM在区分优选与次优回答上的准确性。结果表明,尽管两者在训练集上表现相当,但DPORM在验证集上的泛化能力更弱,尤其在分布外场景下。在五个分布外设置中,DPORM平均准确率下降3%,最大下降达7%。这说明DPORM泛化能力有限,支持在迭代DPO中引入显式奖励模型。

原文摘要 · Abstract (English)

Reinforcement Learning from Human Feedback (RLHF) is an effective approach for aligning language models to human preferences. Central to RLHF is learning a reward function for scoring human preferences. Two main approaches for learning a reward model are 1) training an EXplicit Reward Model (EXRM) as in RLHF, and 2) using an implicit reward learned from preference data through methods such as Direct Preference Optimization (DPO). Prior work has shown that the implicit reward model of DPO (denoted as DPORM) can approximate an EXRM in the limit. DPORM's effectiveness directly implies the optimality of the learned policy, and also has practical implication for LLM alignment methods including iterative DPO. However, it is unclear how well DPORM empirically matches the performance of EXRM. This work studies the accuracy at distinguishing preferred and rejected answers for both DPORM and EXRM. Our findings indicate that even though DPORM fits the training dataset comparably, it generalizes less effectively than EXRM, especially when the validation datasets contain distribution shifts. Across five out-of-distribution settings, DPORM has a mean drop in accuracy of 3% and a maximum drop of 7%. These findings highlight that DPORM has limited generalization ability and substantiates the integration of an explicit reward model in iterative DPO approaches.

强化学习偏好学习奖励模型泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。