arXiv:2606.09735cs.CL2026-06

RLHF让大模型表面中立,实则保留了政治倾向的底层结构。

The Neutral Mask: How RLHF Provides Shallow Alignment while Leaving Partisan Structure Intact in a Large Language Model

  • 通过稀疏自编码器发现,指令模型中原本活跃的政治特征完全失效。
  • 模型输出始终中立,但底层政治倾向结构未被消除,仅切断了其输出路径。
  • 适合关注模型对齐机制脆弱性与潜在滥用风险的研究者阅读。

对齐训练旨在使大语言模型安全且有用,主要手段是基于人类反馈的强化学习(RLHF),通过与‘人类价值观’对齐来塑造部署模型的行为。然而该过程仍不透明:编码了何种价值?谁的价值?如何编码?越来越多证据表明,RLHF仅产生功能性合规,而非深层对齐。本文以Llama 3.1 8B为例,对比基线模型与指令模型在政治倾向上的内部表征,发现RLHF并未消除模型中固有的政治方向结构,而是压缩了政治信号的方差,生成一致平衡的非党派输出。稀疏自编码器分解显示,基线模型中偶发激活的策略编码特征在指令模型中完全失活。特征级控制实验进一步证实其因果断联。因此,RLHF并非消除党派结构,而是切断其到输出的因果路径,形成功能性的政治中立。这种中立并非结构性,底层几何结构仍支持党派操控。例如,通过推断用户政治身份并放大其倾向,可重新激活党派生成。若RLHF本质是断连而非移除价值结构,则此模式可能适用于其他价值领域,意味着对齐模型行为比其输出所暗示的更脆弱。

原文摘要 · Abstract (English)

The ambition behind alignment training is to make large language models safe and useful. The primary mechanism, reinforcement learning from human feedback (RLHF), shapes the behavior of deployed language models by aligning them with ``human values.'' Yet the process is opaque. What values are being encoded; whose values are they; and how does RLHF encode them? A growing body of evidence suggests that RLHF produces only functional compliance rather than deep alignment. We offer a mechanistic case study of this phenomenon for partisan political orientation with a comparison of the internal representations of Llama 3.1 8B before and after RLHF. We show that RLHF does not remove the structured partisan direction in the base model. Instead, it compresses the variance of the partisan signal to generate consistently balanced and non-partisan output. Sparse autoencoder decomposition reveals that policy-encoding features, which activate sporadically in the base model, are completely inactive in the Instruct model. Feature-level steering experiments confirm the causal disconnect. RLHF thus encodes a norm of political neutrality, not by erasing the model's knowledge of partisanship, but by severing the causal pathway from partisan geometry to output generation. Importantly, this neutrality is functional, not structural so that the underlying geometry that enables partisan steering remains intact. The mechanisms that bypass RLHF's guardrails, such as inferring and amplifying a user's partisan identity, reactivate partisan generation. If RLHF operates by disconnecting rather than removing value-laden structure, then the same pattern may hold for other value domains, and the aligned model's behavior may be more fragile than its outputs suggest.

模型对齐政治偏见强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。