小样本下用LLM模拟调查数据,提升三重一致性。
Beyond the Mean: Three-Axis Fidelity for Aligning LLM-Based Survey Simulators from Small Pilot Data

- 分解三轴评估:结构、边际、个体一致性。
- 微调在小样本中表现最优,但子群体差异显著。
- 适合需要真实社会数据模拟的社科研究者。
大语言模型(LLMs)被广泛用于模拟社会调查响应,但其输出存在系统性偏差:边缘分布失真、响应方差校准不足,预测变量与结果间关系弱化。本文以新冠虚假信息调查为例,探讨在少量人类响应样本基础上,是否能通过LLM恢复总体统计特征。提出从三个维度评估恢复能力:结构保真度、边际保真度与个体保真度。对比提示工程、校正与微调三种方法,发现微调在小样本下能较好平衡多维保真度,但不同子样本间的保真水平存在差异,可能影响多元共存对齐。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used to simulate social survey responses, yet their outputs exhibit systematic biases: marginal distributions are skewed, response variance is poorly calibrated, and predictor-outcome relationships are attenuated. We ask a simple question: given a small pilot sample of human responses, can an LLM recover the statistical characteristics of a broader population? We decompose recovery along three axes: structural fidelity, marginal fidelity, and individual fidelity. Using a COVID-19 misinformation survey as a case study, we benchmark three families of approaches: prompting, rectification, and fine-tuning. The findings suggest that fine-tuning on small pilot samples offers a balanced approach for achieving multiple forms of fidelity, but the levels of such fidelity can vary across subsamples, potentially threatening pluralistic alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。