arXiv:2606.17165stat.MEcs.AI2026-06被引 2

用大模型替代人类做A/B测试需满足特定假设,否则结果可能严重失真。

Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference

  • 提出基于代理终点的统计框架,通过校准大模型与人类数据提升因果推断可靠性
  • 实验证明未校准的大模型仅能捕获39%的人类处理效应,校准后显著提升
  • 强调大模型随机性会引入偏差,建议对多次采样取平均以改善效果

越来越多组织和研究者希望用大语言模型(LLMs)替代人类参与A/B测试,以实现更快、更低成本的实验。本文研究在何种条件下,基于LLM结果估计的处理效应可还原真实人类群体的因果效应。若LLM与人类结果分布完全等价,则标准估计器有效,但该假设不现实。为此,本文将代理终点理论拓展至LLMs,证明在代理性和可比性条件下,通过校准可识别平均处理效应,且该条件组合弱于分布等价。提出代理性可验证性检验,并给出有限重叠导致最坏偏差的界。进一步发现LLM固有的随机性会削弱代理性并引入估计偏差与方差,但对每个样本取多轮生成结果的均值可缓解此问题。模拟验证结果,基于Upworthy Research Archive数据的实证显示,原始LLM输出仅恢复39%的人类处理效应,而非参数校准后差距显著缩小。核心结论是:基于LLM的A/B测试正确性依赖假设,而人类实验天然正确;但正是在大模型最有望带来效益的场景下,这些假设最难成立。讨论了模型选择、提示设计与温度设置等变量,长期结果带来的复合挑战,以及如何设计人类预试验进行验证。

原文摘要 · Abstract (English)

Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at lower cost. We study when a treatment effect estimated on LLM outcomes can recover the effect for the human population of interest. Distributional equivalence between LLM and human outcomes would make any standard estimator valid but is unrealistic. We therefore develop a statistical framework that adapts surrogate endpoint theory to LLMs, showing that calibrating LLM outcomes to human outcomes identifies the average treatment effect under surrogacy and comparability conditions that are jointly weaker than distributional equivalence. We present a falsification test for surrogacy and a bound on the worst-case bias from limited overlap between the LLM and human samples. We further show that the stochasticity inherent to LLMs can weaken surrogacy for identification while also introducing bias and variance during estimation, but that using an average over multiple LLM draws per unit as the surrogate mitigates these issues. Simulations validate the results, and an empirical application to the Upworthy Research Archive dataset shows that raw LLM outputs recover only 39% of the human treatment effect while nonparametric calibration closes the gap. A central takeaway is that A/B testing on LLM responses is correct only by assumption, whereas A/B testing on humans is correct by design, and that the required assumptions are hardest to justify precisely where LLMs promise the greatest benefit. We discuss the choice of LLM, prompting, and temperature as design variables, the compounded challenge posed by long-term outcomes, and how to size human pilot studies for validation.

A/B测试大模型评估因果推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。