arXiv:2411.04991cs.AI2024-11被引 46

重新审视偏好建模中布拉德利-特瑞模型的理论基础,提出更优替代方案。

Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

  • 基于深度嵌入的BT模型收敛速率获理论证明
  • 实验证明新方法在12000+设置下表现更优
  • 适合需要高效排序一致性的对齐研究者

布拉德利-特瑞(BT)模型是大语言模型对齐中奖励建模的常用方法。然而,其为何能将成对响应比较转化为奖励值并做出预测仍不清晰,尤其在仅有限提示-响应对被稀疏比较的情况下。本文首先重新审视了BT模型在奖励建模中的基础,基于深度神经网络嵌入建立了其收敛速率的理论依据。尽管理论上成立,但作者认为从下游优化角度看,BT并非必要选择——奖励模型只需通过单调变换保持真实奖励的正确排序即可。文中强调了排序一致性的重要性,并证明了BT具备该性质。因此,提出一种简单直接的上界算法,兼容现成二分类器,作为排序一致的替代目标。为提供实践洞见,作者在超过12,000个实验设置下评估了不同方法,使用6个基础大模型、2个数据集,以及多样化的标注设计(涵盖数量、质量与配对选择差异)。

原文摘要 · Abstract (English)

The Bradley-Terry (BT) model is a common and successful practice in reward modeling for Large Language Model (LLM) alignment. However, it remains unclear why this model -- originally developed for multi-player stochastic game matching -- can be adopted to convert pairwise response comparisons to reward values and make predictions. Especially given the fact that only a limited number of prompt-response pairs are sparsely compared with others. In this paper, we first revisit the foundations of using BT models in reward modeling, and establish the convergence rate of BT reward models based on deep neural networks using embeddings, providing a theoretical foundation for their use. Despite theoretically sound, we argue that the BT model is not a necessary choice from the perspective of downstream optimization. This is because a reward model only needs to preserve the correct ranking predictions through a monotonic transformation of the true reward. We highlight the critical concept of order consistency in reward modeling and demonstrate that the BT model possesses this property. Consequently, we propose a simple and straightforward upper-bound algorithm, compatible with off-the-shelf binary classifiers, as an alternative order-consistent reward modeling objective. To offer practical insights, we empirically evaluate the performance of these different reward modeling approaches across more than 12,000 experimental setups, using $6$ base LLMs, $2$ datasets, and diverse annotation designs that vary in quantity, quality, and pairing choices in preference annotations.

奖励建模排序一致性大模型对齐理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。