用大模型模拟虚拟人群,零数据预测问卷结果
Llms, Virtual Users, and Bias: Predicting Any Survey Question Without Human Data
- 用大模型生成虚拟受访者回答问卷,无需真实人类数据
- 去除了内容过滤的大模型在少数群体上预测更准
- 适合做民意研究的低成本替代方案,但需警惕偏见
大型语言模型(LLMs)为传统调查方法提供了有前景的替代方案,有望提升效率并降低成本。本研究利用LLMs构建虚拟人口来回答调查问题,实现与人类回答相当的预测效果。我们评估了GPT-4o、GPT-3.5、Claude 3.5-Sonnet以及Llama和Mistral系列模型的表现,并与基于世界价值观调查(WVS)人口统计学数据的随机森林算法进行对比。总体而言,LLMs表现具有竞争力,且无需额外训练数据。然而,它们在预测特定宗教和人口群体的回答时表现出偏差,表现不佳。相比之下,使用足够数据训练的随机森林性能更强。研究发现,移除LLMs的内容审查机制可显著提升预测准确性,尤其在代表性不足的人群中,未审查模型表现更优。这些结果凸显了在公众意见研究中应对模型偏见、重新审视审查策略的重要性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) offer a promising alternative to traditional survey methods, potentially enhancing efficiency and reducing costs. In this study, we use LLMs to create virtual populations that answer survey questions, enabling us to predict outcomes comparable to human responses. We evaluate several LLMs-including GPT-4o, GPT-3.5, Claude 3.5-Sonnet, and versions of the Llama and Mistral models-comparing their performance to that of a traditional Random Forests algorithm using demographic data from the World Values Survey (WVS). LLMs demonstrate competitive performance overall, with the significant advantage of requiring no additional training data. However, they exhibit biases when predicting responses for certain religious and population groups, underperforming in these areas. On the other hand, Random Forests demonstrate stronger performance than LLMs when trained with sufficient data. We observe that removing censorship mechanisms from LLMs significantly improves predictive accuracy, particularly for underrepresented demographic segments where censored models struggle. These findings highlight the importance of addressing biases and reconsidering censorship approaches in LLMs to enhance their reliability and fairness in public opinion research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。