用微调大模型模拟全球调查回答分布,提升社会科学研究效率
Specializing Large Language Models to Simulate Survey Response Distributions for Global Populations
- 基于首词概率微调,缩小预测与真实回答分布差距
- 在未见问题、国家和调查上仍优于零样本方法
- 为社会科学研究提供低成本模拟工具,适合政策分析者
大规模调查是社会科学研究与政策制定的重要工具,但实施成本高、耗时长。若能准确模拟群体层面的调查结果,将极大促进研究进展。已有工作尝试用大语言模型(LLMs)通过提示生成人类行为,但本论文首次针对模拟调查回答分布任务对LLMs进行专业化。以两项全球文化调查的国家层面数据为测试集,提出基于首词概率的微调方法,最小化预测与实际回答分布间的差异。实验表明,该方法显著优于其他方法及零样本分类器,即使在未见问题、国家和完全未见过的调查中依然表现更好。尽管最佳模型在未见问题上仍有挑战,结果证明专业化对模拟任务具有明显优势,有望加速实现足够精准的模拟。
原文摘要 · Abstract (English)
Large-scale surveys are essential tools for informing social science research and policy, but running surveys is costly and time-intensive. If we could accurately simulate group-level survey results, this would therefore be very valuable to social science research. Prior work has explored the use of large language models (LLMs) for simulating human behaviors, mostly through prompting. In this paper, we are the first to specialize LLMs for the task of simulating survey response distributions. As a testbed, we use country-level results from two global cultural surveys. We devise a fine-tuning method based on first-token probabilities to minimize divergence between predicted and actual response distributions for a given question. Then, we show that this method substantially outperforms other methods and zero-shot classifiers, even on unseen questions, countries, and a completely unseen survey. While even our best models struggle with the task, especially on unseen questions, our results demonstrate the benefits of specialization for simulation, which may accelerate progress towards sufficiently accurate simulation in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。