arXiv:2502.16761cs.CL2025-02ACL综述被引 78

用大规模调查数据微调大模型,提升对公众意见分布的预测精度。

Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public Opinions

  • 直接基于调查数据结构微调大模型,替代传统提示工程。
  • 在70K子群体-回答对上训练,使模型预测与真实回应差距缩小46%。
  • 适用于未见调查和子群体,助力高效问卷设计。

大型语言模型(LLMs)为公众意见研究带来新机遇,可在调查设计初期提前预测调查回应。以往方法通过描述子人群作为输入提示引导模型,但难以准确预测人类受访者的回应分布。本文提出直接利用调查数据的独特结构对LLM进行微调。为此,我们构建了规模更大的SubPOP数据集,包含3,362个问题和70,000个子群体-回应对,来自权威民意调查。实验表明,在SubPOP上微调显著提升模型预测与真实回应的一致性,相比基线将模型与人类间的差距降低最多达46%,且在未见调查和子群体上表现出强泛化能力。结果表明,基于调查数据的微调可有效提升对多样化真实子群体意见的预测性能,从而实现更高效的调查设计。代码已开源:https://github.com/JosephJeesungSuh/subpop。

原文摘要 · Abstract (English)

Large language models (LLMs) present novel opportunities in public opinion research by predicting survey responses in advance during the early stages of survey design. Prior methods steer LLMs via descriptions of subpopulations as LLMs' input prompt, yet such prompt engineering approaches have struggled to faithfully predict the distribution of survey responses from human subjects. In this work, we propose directly fine-tuning LLMs to predict response distributions by leveraging unique structural characteristics of survey data. To enable fine-tuning, we curate SubPOP, a significantly scaled dataset of 3,362 questions and 70K subpopulation-response pairs from well-established public opinion surveys. We show that fine-tuning on SubPOP greatly improves the match between LLM predictions and human responses across various subpopulations, reducing the LLM-human gap by up to 46% compared to baselines, and achieves strong generalization to unseen surveys and subpopulations. Our findings highlight the potential of survey-based fine-tuning to improve opinion prediction for diverse, real-world subpopulations and therefore enable more efficient survey designs. Our code is available at https://github.com/JosephJeesungSuh/subpop.

大模型微调民意预测调查数据子群体分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。