arXiv:2603.12795cs.CL2026-03被引 1

用稀疏自编码器消除奖励模型对格式的偏好,不重训也能提升公平性。

SteerRM: Debiasing Reward Models via Sparse Autoencoders

  • 通过对比回复识别风格特征,用稀疏自编码器定位偏差源
  • 推理时抑制偏差特征,平均提升硬分数据集准确率7.3分
  • 适用于多种模型架构,揭示浅层共享的格式偏差模式

奖励模型是对齐流程的关键组件,但常偏向表面风格特征,更青睐排版优美的回答而非语义更优者。现有去偏方法多需重新训练或修改结构,而直接激活抑制因表征纠缠导致性能下降。本文提出SteerRM,首个无需训练的去偏方法,利用稀疏自编码器(SAE)干预。SteerRM通过对比成对回复分离风格影响,基于强度-稳定性准则识别与偏差相关的SAE特征,并在推理时进行抑制。在六种奖励模型的RM-Bench测试中,平均提升硬分数据集准确率7.3点,同时保持整体性能。在Gemma基奖励模型及无格式偏差控制实验中亦验证了跨架构与偏类型的一般性。进一步发现,格式相关特征集中于浅层且可跨模型迁移,揭示了架构层面共有的偏差编码模式。结果表明,无需重训的SAE干预可有效缓解奖励模型偏差,为对齐流程提供实用且可解释的解决方案。

原文摘要 · Abstract (English)

Reward models (RMs) are critical components of alignment pipelines, yet they exhibit biases toward superficial stylistic cues, preferring better-presented responses over semantically superior ones. Existing debiasing methods typically require retraining or architectural modifications, while direct activation suppression degrades performance due to representation entanglement. We propose SteerRM, the first training-free method for debiasing reward models using Sparse Autoencoder (SAE)-based interventions. SteerRM isolates stylistic effects using contrastive paired responses, identifies bias-related SAE features with a strength-stability criterion, and suppresses them at inference time. Across six reward models on RM-Bench, SteerRM improves Hard-split accuracy by 7.3 points on average while preserving overall performance. Results on a Gemma-based reward model and a controlled non-format bias further suggest generalization across RM architectures and bias types. We further find that format-related features are concentrated in shallow layers and transfer across models, revealing shared architecture-level bias encoding patterns. These results show that SAE-based interventions can mitigate reward-model biases without retraining, providing a practical and interpretable solution for alignment pipelines.

奖励模型去偏稀疏自编码器对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。