arXiv:2510.13694cs.LG2025-10被引 6

用信息瓶颈原理防奖励黑客,让大模型更稳定对齐人类意图。

Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking

  • 基于信息瓶颈思想过滤无关特征,减少奖励模型误泛化。
  • 发现奖励黑客在隐空间中表现为明显异常点,偏差越大越危险。
  • 提出新指标和正则化方法,可实时检测并抑制奖励过优化。

尽管基于人类反馈的强化学习(RLHF)在对齐语言模型与人类价值观方面取得成功,但奖励黑客——即对奖励函数过度优化——仍是主要挑战。我们识别出两大障碍:(1) 奖励建模中的奖励误泛化,即奖励模型过度拟合与偏好无关的虚假特征;(2) 强化学习优化过程中缺乏合适的正则化,现有逐标记约束常过度限制策略空间。为此,我们提出InfoRM,一种基于信息瓶颈(IB)原则的信息论奖励建模框架,通过过滤偏好无关信息缓解误泛化。进一步观察到,奖励黑客响应在InfoRM的IB隐空间中表现为显著异常点,其偏离SFT诱导分布的程度可用马氏距离衡量。据此,我们引入IBL,一种分布级正则化,惩罚此类偏离,有效拓展优化空间的同时保持对齐。我们证明IBL在IB隐空间中等价于悲观强化学习目标。最后,我们提出马氏异常概率(MOP),一种量化奖励黑客严重性的统计度量,支持合理的超参数调优及在线缓解,如早期停止。在多种LLM和数据集上的大量实验验证了方法的通用性、有效性以及MOP作为诊断工具的可靠性,共同推动了RLHF的发展。

原文摘要 · Abstract (English)

Despite the success of Reinforcement Learning from Human Feedback (RLHF) in aligning language models with human values, reward hacking-or reward over-optimization-remains a major challenge. We identify two key obstacles to its mitigation: (1) reward misgeneralization in reward modeling, where reward models overfit to spurious, preference-irrelevant features; and (2) the lack of suitable regularization during RL optimization, as existing token-level constraints often over-restrict the policy space. To address these issues, we propose InfoRM, an information-theoretic reward modeling framework based on the Information Bottleneck (IB) principle, which filters out preference-irrelevant information to alleviate reward misgeneralization. We further observe that reward-hacked responses manifest as pronounced outliers in InfoRM's IB latent space, measured by Mahalanobis distance from the SFT-induced distribution. Motivated by this, we introduce IBL, a distribution-level regularization that penalizes such deviations, effectively expanding the optimization landscape while maintaining alignment. We prove that IBL is theoretically equivalent to the pessimistic RL objective within the IB latent space. Finally, we present Mahalanobis Outlier Probability (MOP), a statistical metric for quantifying reward hacking severity, enabling principled hyperparameter tuning and online mitigation such as early stopping. Extensive experiments across diverse LLMs and datasets confirm the generality of our findings, the effectiveness of InfoRM and IBL, and the reliability of MOP as a diagnostic tool-collectively advancing the state of RLHF.

强化学习奖励建模对齐异常检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。