发现并操控大模型中的政治立场方向,可精准改变生成内容倾向。
The Amplifying Mirror: Locating and Steering the Partisan Direction inside a Large Language Model

- 通过分析隐藏层激活,定位到模型中决定政治立场的几何轴。
- 在生成过程中操纵该轴,使输出立场反转或虚构权威信息。
- 揭示政治偏见是模型结构特性,非偶然错误,适合政策与安全研究者。
大型语言模型正快速取代搜索引擎成为人与信息的主要接口。与检索已有内容的搜索引擎不同,语言模型生成基于训练中学习到的内部表征的新文本。本文利用190,491条美国国会议员推文作为标注数据,对Llama 3.1 8B Instruct模型的隐藏状态训练线性探测器,发现第18层存在一条单一几何轴,能以AUC 0.945和Cohen's d 1.94区分共和党和民主党文本;通过稀疏自编码器将该轴分解为可解释的政治特征。在生成过程中因果干预该轴(删除或放大其成分),可系统性改变输出,实现立场反转、立场漂移及结构性虚假权威构建。结果表明,政治偏见并非模糊的涌现属性,而是可精确定位与操控的几何特征。它不是需要修补的缺陷,而是模型编码用户信息时的结构性特征。随着语言模型取代搜索引擎成为知识入口,理解这种产品设计及其后果,对于应对从筛选式信息生态向生成式生态转型带来的法律、社会与政治挑战至关重要。
原文摘要 · Abstract (English)
Large language models are rapicly replacing search engines as the primary interface between people and information. Unlike search engines, which retrieve existing content, LLMs generate novel text shaped by internal representations learned during training. Here we show that partisan political identity is encoded in the model's activation space, and that this direction directly shapes generation. Using 190,491 tweets from sitting members of the U.S. Congress as labeled training data, we train linear probes on the hidden states of the Llama 3.1 8B Instruct model. We identify a single geometric axis at layer 18 that separates Republican from Democratic text with an AUC of 0.945 and a Cohen's d of 1.94, and use sparse autoencoders to decompose that axis into interpretable partisan features. Causally intervening along this axis, ablating or amplifying the partisan component mid-generation, produces systematic shifts in the model's output. We witness stance reversals, register shifting, and structured fabrications of authority. Our results demonstrate that partisan bias in language models is not a vague emergent property but a learned geometric feature that can be precisely located and steered. Partisan bias is not a bug to be patched, but a structural property of how these models encode information about their users. As LLMs displace search engines as the interface to knowledge, understanding that product design (and its consequences) will be essential for navigating the legal, social, and political transitions from an information ecosystem that is curated to one that is generated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。