用大模型模拟人类行为做实验,如何确保结果可信?
This human study did not involve human subjects: Validating LLM simulations as behavioral evidence
- 用统计校准法结合真人数据修正模型偏差,提升推断有效性
- 相比纯真人实验,成本更低且精度更高,适合因果推断
- 适合想快速验证假设的社科研究者,但需注意模型泛化能力
越来越多研究使用大语言模型(LLMs)作为合成被试,在社会科学实验中生成低成本、近乎即时的响应。然而,目前缺乏明确指导说明在何种条件下此类模拟可支持对人类行为的有效推断。本文对比了两种获取因果效应有效估计的方法:启发式方法通过提示工程、微调等策略试图使模拟行为与真实人类行为可互换,适用于探索性研究,但缺乏确认性研究所需的统计保证;统计校准则结合辅助真人数据与统计调整,以弥补观察与模拟响应间的差异。在明确假设下,该方法能保持推断有效性,并以更低成本获得更精确的因果效应估计。但两种方法的效果均依赖于大模型对目标人群的逼近程度。本文还指出,若研究者只关注用模型替代真人被试,可能忽略更广阔的潜在机会。
原文摘要 · Abstract (English)
A growing literature uses large language models (LLMs) as synthetic participants to generate cost-effective and nearly instantaneous responses in social science experiments. However, there is limited guidance on when such simulations support valid inference about human behavior. We contrast two strategies for obtaining valid estimates of causal effects and clarify the assumptions under which each is suitable for exploratory versus confirmatory research. Heuristic approaches seek to establish that simulated and observed human behavior are interchangeable through prompt engineering, model fine-tuning, and other repair strategies designed to reduce LLM-induced inaccuracies. While useful for many exploratory tasks, heuristic approaches lack the formal statistical guarantees typically required for confirmatory research. In contrast, statistical calibration combines auxiliary human data with statistical adjustments to account for discrepancies between observed and simulated responses. Under explicit assumptions, statistical calibration preserves validity and provides more precise estimates of causal effects at lower cost than experiments that rely solely on human participants. Yet the potential of both approaches depends on how well LLMs approximate the relevant populations. We consider what opportunities are overlooked when researchers focus myopically on substituting LLMs for human participants in a study.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。