arXiv:2509.09871cs.CLcs.AI2025-09综述被引 2

用AI生成智利民意调查数据,验证其能否真实反映公众意见。

Emulating Public Opinion: A Proof-of-Concept of AI-Generated Synthetic Survey Responses for the Chilean Case

  • 用128个提示-模型-问题组合生成近19万份虚拟回答。
  • 对信任类问题的准确率和F1分数均超0.90,表现优异。
  • 45-59岁群体的虚拟回答最接近真实人群,需警惕偏差。

大型语言模型(LLMs)为调查研究提供了方法创新可能,通过合成受访者模拟人类回答与行为,或可缓解测量与代表性误差。然而,其在恢复总体项目分布方面的可靠性尚不明确,下游应用还可能复制训练数据中的社会刻板印象与偏见。本文以智利概率抽样调查的真实数据为基准,评估了由LLM生成的合成回答的可靠性。具体测试了128个提示-模型-问题组合,生成189,696份合成样本,并在128个问题-子样本对上进行元分析,综合评估准确率、精确率、召回率与F1分数,检验关键人口统计维度上的偏差。测试涵盖OpenAI的GPT系列与o系列推理模型,以及Llama与Qwen模型。结果表明:第一,合成回答在信任类问题上表现极佳(F1分数与准确率均>0.90);第二,GPT-4o、GPT-4o-mini与Llama 4 Maverick表现相当;第三,45-59岁群体的合成回答与真实人群一致性最高。总体而言,基于LLM的合成样本能较好逼近概率样本的回答特征,但存在显著的项目级差异。全面捕捉公众意见的细微差别仍具挑战,需精细校准与额外分布检验以保障算法保真度、降低误差。

原文摘要 · Abstract (English)

Large Language Models (LLMs) offer promising avenues for methodological and applied innovations in survey research by using synthetic respondents to emulate human answers and behaviour, potentially mitigating measurement and representation errors. However, the extent to which LLMs recover aggregate item distributions remains uncertain and downstream applications risk reproducing social stereotypes and biases inherited from training data. We evaluate the reliability of LLM-generated synthetic survey responses against ground-truth human responses from a Chilean public opinion probabilistic survey. Specifically, we benchmark 128 prompt-model-question triplets, generating 189,696 synthetic profiles, and pool performance metrics (i.e., accuracy, precision, recall, and F1-score) in a meta-analysis across 128 question-subsample pairs to test for biases along key sociodemographic dimensions. The evaluation spans OpenAI's GPT family and o-series reasoning models, as well as Llama and Qwen checkpoints. Three results stand out. First, synthetic responses achieve excellent performance on trust items (F1-score and accuracy > 0.90). Second, GPT-4o, GPT-4o-mini and Llama 4 Maverick perform comparably on this task. Third, synthetic-human alignment is highest among respondents aged 45-59. Overall, LLM-based synthetic samples approximate responses from a probabilistic sample, though with substantial item-level heterogeneity. Capturing the full nuance of public opinion remains challenging and requires careful calibration and additional distributional tests to ensure algorithmic fidelity and reduce errors.

AI生成民意调查大模型合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。