通过因果干预抑制奖励模型中的多种伪特征,提升对齐效果。
Debiasing Reward Models via Causally Motivated Inference-Time Intervention

- 识别与偏见特征强相关的神经元,进行逐层信号抑制。
- 在多个基准上降低对冗长等伪特征的敏感性,无性能损失。
- 小模型经少量修改即可媲美大模型,适合资源有限场景。
奖励模型(RMs)在对齐大语言模型(LLMs)与人类偏好中起核心作用。然而,现有方法在推理阶段通常仅关注响应长度这一单一偏见,导致性能权衡。本文提出一种基于因果动机的推理时干预方法,可同时缓解多种类型的偏见。该方法首先识别与预定义偏见属性强相关的神经元激活,再在神经元层级实施信号抑制。在多个RM基准上的实验表明,该方法显著降低了对各类伪特征的敏感性,且未引发性能下降。当用于偏好标注时,使用该方法的2B和7B小规模奖励模型(仅修改不到2%的神经元),使大语言模型实现更好的对齐,其表现接近于当前最先进的70B RM,在AlpacaEval和MT-Bench上达到相当水平。进一步分析显示,偏见信号主要由早期层的神经元编码,揭示了奖励模型内部偏见利用机制。
原文摘要 · Abstract (English)
Reward models (RMs) play a central role in aligning large language models (LLMs) with human preferences. However, RMs are often sensitive to spurious features such as response length. Existing inference-time approaches for mitigating these biases typically focus exclusively on response length, resulting in performance trade-offs. In this paper, we propose causally motivated intervention for mitigating multiple types of biases in RMs at inference time. Our method first identifies neurons whose activations are strongly correlated with predefined bias attributes, and applies neuron-level intervention that suppresses these signals. We evaluate our method on RM benchmarks and observe reductions in sensitivity to spurious features across diverse bias types, without inducing performance trade-offs. Moreover, when used for preference annotation, small RMs (2B and 7B) with our method, which edits less than 2% of all the neurons in RMs, enable LLMs to improve alignment, achieving performance comparable to that of a state-of-the-art 70B RM on AlpacaEval and MT-Bench. Further analysis reveals that bias signals are primarily encoded by neurons in early layers, shedding light on the internal mechanisms of bias exploitation in RMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。