arXiv:2507.00665cs.CLcs.AI2025-07被引 5

用稀疏自编码器解析奖励模型,提升安全对齐的可解释性与可控性

SAFER: Probing Safety in Reward Models with Sparse Autoencoder

  • 通过稀疏自编码器提取奖励模型中的可解释特征
  • 仅修改少量数据即可精准削弱或增强安全对齐效果
  • 适合关注大模型对齐安全、可解释性的研究者与工程师

强化学习人类反馈(RLHF)是使大语言模型对齐人类价值观的关键范式,但其核心奖励模型仍高度不透明。本文提出稀疏自编码器增强奖励模型(SAFER)框架,通过机械解析方法揭示奖励模型激活中的可解释特征,洞察与安全相关的决策机制。我们基于安全偏好数据集,利用选择与拒绝响应间激活差异量化特征显著性,并据此设计针对性的数据污染与去噪策略。实验表明,SAFER仅需少量数据修改即可精确降低或提升安全对齐程度,且不损害通用对话性能。该方法为高风险大模型对齐任务中的解释、审计与优化提供了新路径。代码已开源:https://github.com/xzy-101/SAFER-code。本文涉及奖励模型安全性话题,可能包含潜在风险或非安全案例。

原文摘要 · Abstract (English)

Reinforcement learning from human feedback (RLHF) is a key paradigm for aligning large language models (LLMs) with human values, yet the reward models at its core remain largely opaque. In this work, we present Sparse Autoencoder For Enhanced Reward model (\textbf{SAFER}), a novel framework for interpreting and improving reward models through mechanistic analysis. Leveraging Sparse Autoencoders (SAEs), we uncover human-interpretable features in reward model activations, enabling insight into safety-relevant decision-making. We apply SAFER to safety-oriented preference datasets and quantify the salience of individual features by activation differences between chosen and rejected responses. Using these feature-level signals, we design targeted data poisoning and denoising strategies. Experiments show that SAFER can precisely degrade or enhance safety alignment with minimal data modification, without sacrificing general chat performance. Our approach contributes to interpreting, auditing and refining reward models in high-stakes LLM alignment tasks. Our codes are available at https://github.com/xzy-101/SAFER-code. \textit{This paper discusses topics related to reward model safety and may include discussions or examples that highlight potential risks or unsafe outcomes.}

奖励模型安全对齐可解释性稀疏编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。