实验数据虽能发现观察研究的偏差,但无法验证其因果估计的准确性。
The Hardness of Validating Observational Studies with Experimental Data
- 用实验数据检验观察数据的因果估计是否偏误
- 无额外假设下,实验数据无法消除观察数据的系统性偏差
- 提出基于高斯过程的新方法,构建可信的因果效应区间
观察数据虽易获取且量大,但常因未观测混杂变量导致因果效应估计偏倚。近期研究尝试通过补充小规模实验数据(如随机对照试验)来纠正偏倚。本文证明了一个基础性定理:尽管实验数据可用于‘否定’观察数据的因果估计,但一般情况下无法‘验证’其正确性。在‘不可能推断’框架下,我们指出,若不假设校正函数的平滑性,实验数据无法消除偏倚。为此,我们提出一种基于高斯过程的新方法,构建在实验数据支持范围内外均具有高置信度的因果效应区间。在模拟与半合成数据上验证了方法有效性,并开源代码。
原文摘要 · Abstract (English)
Observational data is often readily available in large quantities, but can lead to biased causal effect estimates due to the presence of unobserved confounding. Recent works attempt to remove this bias by supplementing observational data with experimental data, which, when available, is typically on a smaller scale due to the time and cost involved in running a randomised controlled trial. In this work, we prove a theorem that places fundamental limits on this ``best of both worlds'' approach. Using the framework of impossible inference, we show that although it is possible to use experimental data to \emph{falsify} causal effect estimates from observational data, in general it is not possible to \emph{validate} such estimates. Our theorem proves that while experimental data can be used to detect bias in observational studies, without additional assumptions on the smoothness of the correction function, it can not be used to remove it. We provide a practical example of such an assumption, developing a novel Gaussian Process based approach to construct intervals which contain the true treatment effect with high probability, both inside and outside of the support of the experimental data. We demonstrate our methodology on both simulated and semi-synthetic datasets and make the \href{https://github.com/Jakefawkes/Obs_and_exp_data}{code available}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。