大模型能从聊天记录推断政治立场,隐私风险远超想象。
LLMs Can Infer Political Alignment from Online Conversations
- 用社交媒体文本训练大模型,通过非政治词汇推断立场
- 准确率显著高于传统机器学习,多轮推理后效果更优
- 适合关注AI隐私风险的研究者与政策制定者
由于身份、文化与政治态度等特征之间的相关性,看似无害的偏好(如追某乐队或使用特定俚语)可能暴露个人私密属性。当海量公开社交数据与先进计算方法结合时,这一现象带来根本性隐私威胁。随着线上数据暴露增加和人工智能快速发展,理解大语言模型(LLMs)利用此类信息的能力至关重要。本文基于DebateOrg与Reddit上的在线讨论,证明LLMs可可靠推断隐藏的政治立场,显著优于传统机器学习模型。预测准确率随文本层面推理的聚合及更多政治相关领域数据的引入而提升。研究发现,LLMs依赖高度预测性强但非显式政治性的词汇。结果凸显了大模型在利用社会文化关联方面的强大能力及其潜在风险。
原文摘要 · Abstract (English)
Due to the correlational structure in our traits such as identities, cultures, and political attitudes, seemingly innocuous preferences like following a band or using a specific slang can reveal private traits. This possibility, especially when combined with massive, public social data and advanced computational methods, poses a fundamental privacy risk. As our data exposure online and the rapid advancement of AI are increasing the risk of misuse, it is critical to understand the capacity of large language models (LLMs) to exploit such potential. Here, using online discussions on DebateOrg and Reddit, we show that LLMs can reliably infer hidden political alignment, significantly outperforming traditional machine learning models. Prediction accuracy further improves as we aggregate multiple text-level inferences into a user-level prediction, and as we use more politics-adjacent domains. We demonstrate that LLMs leverage words that are highly predictive of political alignment while not being explicitly political. Our findings underscore the capacity and risks of LLMs for exploiting socio-cultural correlates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。