用大模型预测+实验数据,让随机实验更省样本、更准。
Efficient Randomized Experiments Using Foundation Models
- 融合多个大模型预测与真实实验数据,提升估计精度。
- 最多可减少20%样本量,达到相同统计精度。
- 即使模型预测偏差大,结论依然可信,适合临床/政策评估。
随机实验是评估干预效果的首选方法,但成本高且估计不确定性大。而基于基础模型的模拟实验成本低,可能实现更高统计精度。然而,若模型无法准确预测干预响应,推断结果将无效。本文提出一种新方法,将多个基础模型的预测与实验数据结合,在保持有效统计推断的前提下,获得一致且渐近正态的估计量,其渐近方差不高于仅使用实验数据的标准估计量。重要的是,该性质在模型预测任意有偏时仍成立。多个随机实验的实证结果表明,该方法可显著提升精度,相当于减少最多20%的样本量即可达到与纯实验方法相同的精度。
原文摘要 · Abstract (English)
Randomized experiments are the preferred approach for evaluating the effects of interventions, but they are costly and often yield estimates with substantial uncertainty. On the other hand, in silico experiments leveraging foundation models offer a cost-effective alternative that can potentially attain higher statistical precision. However, the benefits of in silico experiments come with a significant risk: statistical inferences are not valid if the models fail to accurately predict experimental responses to interventions. In this paper, we propose a novel approach that integrates the predictions from multiple foundation models with experimental data while preserving valid statistical inference. Our estimator is consistent and asymptotically normal, with asymptotic variance no larger than the standard estimator based on experimental data alone. Importantly, these statistical properties hold even when model predictions are arbitrarily biased. Empirical results across several randomized experiments show that our estimator offers substantial precision gains, equivalent to a reduction of up to 20% in the sample size needed to match the same precision as the standard estimator based on experimental data alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。