大模型内部存在线性政治立场表示,可被探测与调控。
Linear Representations of Political Perspective Emerge in Large Language Models
- 用注意力头激活值线性预测美国议员意识形态得分。
- 中间层注意力头最能预测政治立场,且可跨文本类型泛化。
- 可通过干预实现生成内容立场的可控调节,适合对齐研究者。
大型语言模型(LLMs)能生成体现不同主观政治视角的文本。本文研究模型如何反映美国政治中更自由或更保守的立场。我们发现,三款开源Transformer模型(Llama-2-7b-chat、Mistral-7b-instruct、Vicuna-7b)在激活空间中存在政治立场的线性表示,相似立场在空间中距离更近。通过提示模型生成不同美国议员视角的文本,我们识别出能线性预测这些议员DW-NOMINATE得分(政治意识形态标准度量)的注意力头。这些高度预测性的头主要位于中间层,常被认为编码高层概念。仅用议员意识形态训练的探测器,也能预测新闻媒体偏斜度。该线性探针可用于可视化、解释和监控模型生成过程中隐含的政治立场。最后,通过线性干预注意力头,可引导输出趋向更自由或更保守立场。结果表明,大模型具备美国政治意识形态的高层线性表示,结合机械可解释性进展,可实现立场的识别、监控与操控。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated the ability to generate text that realistically reflects a range of different subjective human perspectives. This paper studies how LLMs are seemingly able to reflect more liberal versus more conservative viewpoints among other political perspectives in American politics. We show that LLMs possess linear representations of political perspectives within activation space, wherein more similar perspectives are represented closer together. To do so, we probe the attention heads across the layers of three open transformer-based LLMs (Llama-2-7b-chat, Mistral-7b-instruct, Vicuna-7b). We first prompt models to generate text from the perspectives of different U.S. lawmakers. We then identify sets of attention heads whose activations linearly predict those lawmakers' DW-NOMINATE scores, a widely-used and validated measure of political ideology. We find that highly predictive heads are primarily located in the middle layers, often speculated to encode high-level concepts and tasks. Using probes only trained to predict lawmakers' ideology, we then show that the same probes can predict measures of news outlets' slant from the activations of models prompted to simulate text from those news outlets. These linear probes allow us to visualize, interpret, and monitor ideological stances implicitly adopted by an LLM as it generates open-ended responses. Finally, we demonstrate that by applying linear interventions to these attention heads, we can steer the model outputs toward a more liberal or conservative stance. Overall, our research suggests that LLMs possess a high-level linear representation of American political ideology and that by leveraging recent advances in mechanistic interpretability, we can identify, monitor, and steer the subjective perspective underlying generated text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。