在小安慰剂组场景下,用高斯过程提升因果效应估计的置信区间校准度。
Calibrated Inference for the Conditional Average Treatment Effect in the Few-Placebo Regime via Gaussian Processes
- 用高斯过程建模各组结果曲面,让稀缺组的不确定性直接进入后验分布
- 在合成与半合成数据上实现准确覆盖,而传统方法如X-Learner和因果森林均欠覆盖
- 适合小样本因果推断场景,尤其关注置信区间可靠性研究者
估计干预对个体的平均处理效应(CATE)在医学、经济与政策决策中日益重要,其价值取决于是否配有校准的不确定性区间。本文研究少数安慰剂组情形——即一治疗组远小于另一组,常见于非均衡分配试验和小规模保留A/B测试。标准估计器为X-Learner,自然做法是将其第二阶段贝叶斯化以获得可信区间。我们发现此类区间存在欠覆盖:实际包含真实效应的概率低于名义水平。根源在于X-Learner回归目标继承了对小组拟合的扰动模型偏差,导致后验中心偏移。尽管使用正交双重稳健得分是标准修复手段,但在该设定下仍不可靠,因重叠有限导致估计量或高度变异,或稳定后再度出现偏差。二者皆反映更广泛现象:将独立估计的方差附加给难以学习的点估计,但未捕捉点估计的偏差。为此我们提出GP-CATE,通过高斯过程分别建模每组结果表面,使稀缺组的不确定性直接融入后验。在合成与半合成基准测试中,GP-CATE实现了校准覆盖,而对比方法(包括因果森林和BART)未能做到,代价是数据不信息时区间适当变宽。
原文摘要 · Abstract (English)
Estimating how much an intervention helps a given individual the conditional average treatment effect (CATE) is increasingly central to decision-making in medicine, economics, and policy, where an estimate is most useful when accompanied by a calibrated uncertainty interval. We study the few-placebo regime, in which one treatment arm is much smaller than the other, as arises in unequal-allocation trials and small-holdout $A/B$ tests. The standard estimator in this setting is the X-Learner, and a natural way to obtain credible intervals is to make its second stage Bayesian. We show that these intervals under-cover: they contain the true effect less often than their nominal level. We trace this to a structural cause the X-Learner's regression target inherits the bias of a nuisance model fitted to the small arm, so the posterior is centered away from the true effect and we find that the standard remedy, regressing an orthogonal doubly-robust score, is also unreliable here, since the regime's limited overlap leaves the estimator either highly variable or, once stabilized, biased once more. Both consequences reflect a pattern that extends beyond causal inference: a separately estimated variance is attached to a point estimate of a hard-to-learn quantity, and the point estimate's bias is not captured by that variance. We propose GP-CATE, which models each arm's outcome surface with a Gaussian process, so the scarce arm's uncertainty enters the posterior directly rather than as an unmodelled bias. Across synthetic and semi-synthetic benchmarks, GP-CATE attains calibrated coverage where the estimators we compare against including Causal Forest and BART do not, at the cost of intervals that are appropriately wide when the data are uninformative.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。