LLM模拟实验可能因用户属性漂移导致结果偏差,需用负向控制检测并缓解。
The Illusion of Intervention: Your LLM-Simulated Experiment is an Observational Study

- 用负向控制变量检测干预下用户隐变量的分布偏移
- 干预引发的用户漂移会使效果估计被高估或低估
- 通过补充相关混淆因子可显著降低偏差,适合评估类研究
大型语言模型(LLMs)在模拟人类行为方面具有潜力,能以可扩展方式研究干预响应。然而,由于LLMs主要基于观察数据训练,其在模拟实验中对干预的响应可能导致隐含用户属性的非预期变化,引发用户漂移——即不同处理组间的模拟人群分布差异,从而扭曲效应估计。本文形式化了由用户漂移引起的混杂或选择偏差,并表明干预相关的属性变化可能放大或缩小观测到的响应差异。为诊断混杂,提出使用负向控制结果——在干预下应保持不变的变量——来识别各干预条件下分布的偏移,提供用户漂移的证据。为缓解漂移,研究通过引入额外的混淆因子调整角色设定,发现针对具体场景的相关混淆因子可显著减少调查式与多轮交互评估中的偏差。
原文摘要 · Abstract (English)
Large language models (LLMs) show potential as simulators of human behavior, offering a scalable way to study responses to interventions. However, because LLMs are trained largely on observational data, interventions in experiments with LLM-simulated synthetic users can induce unintended shifts in latent user attributes, causing user drift where the implicit simulated population differs across treatment conditions, potentially distorting effect estimates. We formalize the confounding or selection bias that can arise due to user drift and show how intervention-dependent shifts can inflate or attenuate observed differences in user responses under intervention. To diagnose confounding, we propose using negative control outcomes--attributes that should remain invariant under intervention--to identify distribution shifts across intervention conditions, providing evidence of user drift. To mitigate drift, we study adjusting the persona specification by eliciting additional confounders, finding that targeted, setting-relevant confounders can substantially reduce bias across survey-style and multi-turn agent evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。