用大模型生成虚拟问卷回答,评估其社会人口特征模拟能力
Large Language Models as Virtual Survey Respondents: Evaluating Sociodemographic Response Generation
- 设计两种任务:补全缺失属性与完整生成合成数据
- 在11个真实数据集上测试,发现模型表现具一致性但存在结构输出错误
- 提供可复现的基准,适合研究社会调查自动化的人参考
问卷调查是社会科学和公共政策的基础,但传统方法成本高、耗时长且规模有限。尽管已有研究尝试使用大语言模型(LLMs)作为虚拟受访者,但多数工作局限于单一领域或缺乏统一评估框架。为此,我们提出两种互补任务:部分属性模拟(PAS),即根据不完整信息预测缺失属性;全属性模拟(FAS),即在零样本与上下文增强条件下生成完整合成数据集。二者均作为诊断与探索工具,而非人类数据收集的替代。我们构建了涵盖4个社会学领域的11个真实公共数据集的基准LMM-S^3,评估GPT-3.5/4 Turbo与LLaMA 3.0/3.1-8B在零样本与少样本设置下的表现。结果表明模型家族间存在一致性能趋势,暴露出结构化输出生成的失败模式,并揭示上下文与提示设计对模拟保真度的影响。代码与数据集已开源。
原文摘要 · Abstract (English)
Questionnaire-based surveys are foundational to social science research and public policymaking, yet traditional survey methods remain costly, time-consuming, and often limited in scale. Although prior work has explored large language models (LLMs) as virtual survey respondents, existing studies often address narrow task settings, focus on single sociological domains, or lack a unified evaluation framework that enables systematic comparison across diverse datasets and models. To address these gaps, we introduce two complementary task abstractions: Partial Attribute Simulation (PAS), where LLMs predict missing attributes from incomplete respondent profiles, and Full Attribute Simulation (FAS), where LLMs generate complete synthetic datasets under zero-context and context-enhanced conditions. Both are framed as diagnostic and exploratory tools rather than replacements for human data collection. We curate LLM-S^3 (Large Language Model-based Sociodemographic Survey Simulation), a benchmark spanning 11 real-world public datasets across four sociological domains, and evaluate GPT-3.5/4 Turbo and LLaMA 3.0/3.1-8B under zero-shot and few-shot settings. Our evaluation reveals consistent performance trends across model families, highlights failure modes in structured output generation, and demonstrates how context and prompt design affect simulation fidelity. Our code and dataset are available at: https://github.com/dart-lab-research/LLM-S-Cube-Benchmark
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。