首次大规模追踪12个大模型在2024美国大选期间的表现
Large-Scale, Longitudinal Study of Large Language Models During the 2024 US Election Season
- 用近万条每日提问追踪模型行为变化
- 发现模型对政治倾向引导敏感,且隐含选举预测
- 适合关注AI与政治互动的研究者和政策制定者
2024年美国总统大选是美国主流使用大型语言模型(LLMs)后的首次重大选举。借鉴此前媒体变革(尤其是社交媒体在定向传播与政治极化中的作用)的经验,当前亟需探讨LLMs如何影响信息生态与政治对话。尽管平台已宣布部分选举防护措施,但其实际效果尚不明确。为此,我们开展了一项大规模、纵向研究,对12个模型进行了从7月到11月的近日常态化调查,共提出超过12,000个结构化问题。研究系统性地改变内容与格式,形成丰富数据集,可用于分析模型随时间的行为变化(如模型更新后)、对引导的敏感性、指令响应能力及与选举相关的知识与‘信念’。后半部分进行四项分析:(i)研究选举季中模型行为的纵向变化;(ii)展示选举相关回答对人口统计引导的敏感性;(iii)探究模型对候选人特质的‘信念’;(iv)揭示模型对选举结果的隐含预测。为促进未来在选举语境下对LLMs的评估,我们详细阐述了从问题生成到查询流程及第三方工具的完整方法,并公开发布数据集于https://huggingface.co/datasets/sarahcen/llm-election-data-2024。
原文摘要 · Abstract (English)
The 2024 US presidential election is the first major contest to occur in the US since the popularization of large language models (LLMs). Building on lessons from earlier shifts in media (most notably social media's well studied role in targeted messaging and political polarization) this moment raises urgent questions about how LLMs may shape the information ecosystem and influence political discourse. While platforms have announced some election safeguards, how well they work in practice remains unclear. Against this backdrop, we conduct a large-scale, longitudinal study of 12 models, queried using a structured survey with over 12,000 questions on a near-daily cadence from July through November 2024. Our design systematically varies content and format, resulting in a rich dataset that enables analyses of the models' behavior over time (e.g., across model updates), sensitivity to steering, responsiveness to instructions, and election-related knowledge and "beliefs." In the latter half of our work, we perform four analyses of the dataset that (i) study the longitudinal variation of model behavior during election season, (ii) illustrate the sensitivity of election-related responses to demographic steering, (iii) interrogate the models' beliefs about candidates' attributes, and (iv) reveal the models' implicit predictions of the election outcome. To facilitate future evaluations of LLMs in electoral contexts, we detail our methodology, from question generation to the querying pipeline and third-party tooling. We also publicly release our dataset at https://huggingface.co/datasets/sarahcen/llm-election-data-2024
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。