大模型把政治偏见当知识输出,可能加剧思想单一化。
From Echo Chambers to Epistemic Monoculture: Large Language Models Present Temporally Contingent Partisan Alignments as Knowledge

- 发现大模型内部存在可定位的政治立场方向。
- 训练截止时间影响模型对政治议题的表述方式。
- 适合关注AI政治偏见与信息生态的读者。
大型语言模型正成为公民获取政治信息的新接口,常被比作‘更好的谷歌’。然而这一类比在民主政治中并不恰当:搜索引擎检索人类撰写的内容,而语言模型生成新文本,必然嵌入隐性框架。由于知识传递涉及框架选择,生成式系统无法作为‘所有人类知识’的中立通道。我们发现,政治身份在Llama 3.1 8B模型中以可定位的几何方向编码存在,且对齐训练仅掩盖而非消除该结构。利用模型训练截止于2024年的时机——恰逢美国政治重大重组及健康政策剧变——我们开展操控实验,结果表明模型将随时间变化的政治立场呈现为‘知识’,缺乏区分事实与观点的能力。这使信息环境从回音室演变为认知单一化,大模型虽声称总结‘所有人类知识’,实则放大了训练数据中的文化与党派分裂。
原文摘要 · Abstract (English)
Large language models (LLMs) are rapidly becoming an interface between citizens and political information. They are often regarded as "a better Google." While this analogy might work for some instances, it is unintuitively problematic for democratic politics. A search engine retrieves human-authored documents, while a language model generates novel text that necessarily embeds invisible framing decisions. Because conveying knowledge involves framing, a system that generates answers cannot serve as a neutral conduit to "all human knowledge." Instead, these systems are becoming a new kind of political intermediary. Mechanistic evidence shows that partisan identity is encoded as a locatable geometric direction inside the Llama 3.1 8B model, and that alignment training masks rather than removes this structure. Building on that evidence, we present steering experiments that exploit a model's training cutoff in 2024. This cutpoint auspiciously falls just before a dramatic realignment in American politics marked by the second Trump administration and the MAHA transformation of health politics, providing us with a natural experiment. We find that the model presents temporally contingent partisan alignments as knowledge, with no mechanism for distinguishing fact from opinion. This reality moves the information environment beyond the echo chamber toward an epistemic monoculture where language models, purporting to summarize "all human knowledge" are, in actuality, simply magnifying the cultural and partisan divides inherent in their training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。