arXiv:2602.10623cs.LGcs.AI2026-02中稿 · ICML被引 4

用贝叶斯非负建模提升大模型对齐的可靠性,减少奖励黑客问题。

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling

  • 基于贝叶斯非负因子分析构建双层奖励生成机制,分离实例特征与全局偏见。
  • 在多个数据集上显著降低奖励过优化,分布外性能提升23%以上。
  • 适合关注大模型对齐鲁棒性与可解释性的研究者和工程师。

从人类偏好学习的奖励模型是通过强化学习进行人类反馈对齐大语言模型的核心,但常因标注噪声和系统偏差(如响应长度或风格)而易受奖励黑客攻击。本文提出贝叶斯非负奖励模型(BNRM),将非负因子分析融入布拉德利-特里偏好模型。BNRM通过稀疏、非负的潜在因子生成过程,在两个互补层面表示奖励:实例特定的潜在变量生成解耦的奖励表示,而全局潜在因子的稀疏性作为隐式去偏机制,抑制虚假相关性。这种‘解耦-去偏’结构实现了鲁棒的不确定性感知奖励学习。为使BNRM适用于现代大语言模型,我们设计了一种基于深度模型表征的摊销变分推断网络,支持高效端到端训练。大量实验证明,相比强基线,BNRM显著缓解奖励过优化,提升分布外鲁棒性,并生成更具可解释性的奖励分解。

原文摘要 · Abstract (English)

Reward models learned from human preferences are central to aligning large language models (LLMs) via reinforcement learning from human feedback, yet they are often vulnerable to reward hacking due to noisy annotations and systematic biases such as response length or style. We propose Bayesian Non-Negative Reward Model (BNRM), a principled reward modeling framework that integrates non-negative factor analysis into Bradley-Terry (BT) preference model. BNRM represents rewards through a sparse, non-negative latent factor generative process that operates at two complementary levels: instance-specific latent variables induce disentangled reward representations, while sparsity over global latent factors acts as an implicit debiasing mechanism that suppresses spurious correlations. Together, this disentanglement-then-debiasing structure enables robust uncertainty-aware reward learning. To scale BNRM to modern LLMs, we develop an amortized variational inference network conditioned on deep model representations, allowing efficient end-to-end training. Extensive empirical results demonstrate that BNRM substantially mitigates reward over-optimization, improves robustness under distribution shifts, and yields more interpretable reward decompositions than strong baselines.

大模型对齐奖励建模贝叶斯方法鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。