arXiv:2507.07375cs.LGcs.CL2025-07被引 7

融合博弈论与多目标回归,提升大模型奖励模型的鲁棒性与评分能力。

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary

  • 联合训练博弈论与多目标回归模型,共享嵌入空间增强表达。
  • 在分布外场景下,7B模型性能超越70B基线,抗奖励漏洞能力显著提升。
  • 适合追求高鲁棒性奖励模型的RLHF研究者与工业应用团队。

基于人类偏好数据训练的奖励模型在强化学习中的人类反馈(RLHF)框架下,有效对齐大语言模型与人类意图。然而,RLHF仍易受奖励劫持影响,即策略利用奖励函数缺陷而非真正学习预期行为。尽管已有大量工作缓解此问题,但主要聚焦于同分布场景。本文实证表明,现有方法在更困难的分布外(OOD)设置中表现不佳。我们进一步证明,引入细粒度多属性评分可应对该挑战。但高质量数据稀缺常导致多目标奖励函数性能弱化,成为瓶颈。为此,我们提出统一框架,联合训练基于博弈论(BT)的单目标与基于回归的多目标奖励函数,共享嵌入空间。理论上建立了BT损失与回归目标的联系,揭示其互补优势:回归任务增强单目标模型在复杂OOD场景下的抗奖励劫持能力;而基于BT的训练提升多目标模型的评分性能,使7B模型优于70B基线。大量实验显示,该框架显著提升奖励模型的鲁棒性与评分性能。

原文摘要 · Abstract (English)

Reward models trained on human preference data have demonstrated strong effectiveness in aligning Large Language Models (LLMs) with human intent under the framework of Reinforcement Learning from Human Feedback (RLHF). However, RLHF remains vulnerable to reward hacking, where the policy exploits imperfections in the reward function rather than genuinely learning the intended behavior. Although significant efforts have been made to mitigate reward hacking, they predominantly focus on and evaluate in-distribution scenarios, where the training and testing data for the reward model share the same distribution. In this paper, we empirically show that state-of-the-art methods struggle in more challenging out-of-distribution (OOD) settings. We further demonstrate that incorporating fine-grained multi-attribute scores helps address this challenge. However, the limited availability of high-quality data often leads to weak performance of multi-objective reward functions, which can negatively impact overall performance and become the bottleneck. To address this issue, we propose a unified reward modeling framework that jointly trains Bradley--Terry (BT) single-objective and multi-objective regression-based reward functions using a shared embedding space. We theoretically establish a connection between the BT loss and the regression objective and highlight their complementary benefits. Specifically, the regression task enhances the single-objective reward function's ability to mitigate reward hacking in challenging OOD settings, while BT-based training improves the scoring capability of the multi-objective reward function, enabling a 7B model to outperform a 70B baseline. Extensive experimental results demonstrate that our framework significantly improves both the robustness and the scoring performance of reward models.

奖励建模RLHFOOD鲁棒性多目标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。