arXiv:2604.02458cs.CYcs.AI2026-04

大模型模拟人类反应时,真实度高不等于能准确估计干预效果。

Statistical realism is not evidence that LLMs can estimate treatment effects in social science experiments

  • 用跨国家实验测试大模型模拟数据的统计真实度与处理效应估计准确性。
  • 真实度与处理效应准确率相关性弱,优化真实度反而可能降低准确率。
  • 行为类结果误差更大,适合做政策模拟的人需警惕模型外推风险。

大语言模型(LLMs)被广泛用于模拟人类反应并估计社会科学研究中干预措施的效果,尤其在真实实验成本高昂或不可行时。目前常以统计真实度(即模拟响应是否复现真实人类响应的统计特征)作为评估标准,但其能否预测处理效应估计准确性尚不清楚。本文通过一项包含59,508名参与者、来自62个国家的跨国实验,使用三种LLMs联合测量了同一组模拟响应的统计真实度与处理效应准确率。结果显示两者相关性微弱,且在模型、提示词和目标人群选择中优化真实度反而会恶化处理效应估计。该现象在另外两项涵盖12国和27国、共20,785名参与者的跨国实验中重复验证。这种偏差源于两类评估指标不同的误差结构,尤其在行为类结果中更为显著,表明模型可能从态度模式外推行为效应。由于此类误差在部署中难以察觉,可能导致决策误导。本文提出一个诊断框架,用于评估LLM生成的合成数据,并讨论在不同实验基准可用性下如何开展处理效应验证。模拟响应与模拟处理效应是两个独立的估计目标,一方证据不能证明另一方有效。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to simulate human responses and estimate treatment effect of interventions when real-world experiments are costly or infeasible. The treatment-effect estimates are often evaluated using statistical realism, the degree to which simulated responses reproduce properties of observed human responses, although whether realism predicts treatment-effect accuracy remains unknown. Here we test this proxy relationship by jointly measuring statistical realism and treatment-effect accuracy on the same simulated responses in a cross-national experiment with 59,508 participants from 62 countries using three LLMs. The correlation between statistical realism and treatment-effect accuracy is weak, and optimizing for statistical realism can even worsen treatment-effect accuracy when selecting models, prompts, and target populations. The pattern replicates in two additional cross-national experiments spanning 12 and 27 countries with 20,785 participants. The divergence between the two reflects distinct error structures and is larger for behavioral outcomes, where models appear to extrapolate behavioral effects from attitudinal patterns. Because this divergence may remain hidden in deployment, errors can propagate into simulation-informed decisions. We introduce a diagnostic framework for LLM-generated synthetic data and discuss how treatment-effect validation should proceed under varying availability of experimental benchmarks. Simulated responses and simulated treatment effects are distinct estimation targets, and evidence for one does not certify the other.

大模型因果推断仿真评估社会实验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。