arXiv:2608.15871cs.CYcs.CL2026-08

用大模型从人口画像推断投票行为,精度接近真实选举结果。

Large Language Models as Implicit Sociological Models: Reconstructing Voting Behaviour from Sociodemographic Profiles

论文配图:Large Language Models as Implicit Sociological Models: Reconstructing Voting Behaviour from Sociodemographic Profiles
图 1 · 摘自论文原文
  • 用人口特征条件化大模型,生成投票概率并软投票聚合。
  • 在捷克2021年议会选举中误差低,复现了政治阵营与社会梯度。
  • 提供新方法工具,但强调其非预测性且有伦理边界。

大规模语言模型(LLMs)在互联网语料上训练后,编码了关于社会身份、态度和政治行为的广泛统计规律。本文提出并评估了一种方法框架,利用这些潜在表征,从个体层面的人口统计学资料重构集体投票行为。我们通过将大模型以人口描述为条件,提取出投票概率和政党偏好,并采用软投票程序聚合个体输出。以2021年捷克议会选举为验证案例,结果表明当代大模型能以低均方绝对误差重现官方选举结果,恢复已知的政治派系结构,并与独立确立的社会人口梯度一致。本研究贡献在于方法论而非预测:展示了如何系统性地探查大模型作为社会现实的压缩表征,为计算社会科学提供一种新型探索工具,同时明确界定其认识论与伦理边界。

原文摘要 · Abstract (English)

Large language models (LLMs) trained on large-scale internet corpora encode extensive statistical regularities about social identities, attitudes, and political behaviour. This paper introduces and evaluates a methodological framework that leverages these latent representations to reconstruct aggregate voting behaviour from individual-level sociodemographic profiles. We operationalize LLMs as implicit sociological models by conditioning them on demographic descriptions, eliciting probabilistic turnout and party preferences, and aggregating individual outputs via a soft voting procedure. Using the 2021 Czech parliamentary election as a validation case, we demonstrate that contemporary LLMs reproduce official election outcomes with low mean absolute error, recover known political bloc structures, and align with independently established sociodemographic gradients. The contribution of this work is methodological rather than predictive: we show how LLMs can be systematically interrogated as compressed representations of social reality, offering a novel exploratory instrument for computational social science while clearly delineating its epistemic and ethical limits.

大模型社会计算投票行为人口统计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。