利用未标注数据提升因果推断效率,降低估计方差。
Prediction-Powered Causal Inference by Automatic Debiased Machine Learning and Semi-Supervised Riesz Regression
- 结合去偏机器学习与半监督瑞斯回归,构建高效因果估计器。
- 理论证明可实现比仅用标注数据更低的渐近方差。
- 适合关注高精度因果分析的研究者,尤其有辅助变量可用时。
本研究探讨在半监督设定下因果与结构参数的半参数高效估计。设定中除带标签的观测数据(包含结果与协变量)外,还存在未标注的辅助协变量。目标是构造因果与结构参数估计量,使其渐近方差低于仅使用标注数据所得估计量。我们称此框架为预测驱动的因果推断(PPCI)。首先推导出有效影响函数与效率界,表明利用辅助协变量可获得比仅用标注数据更低的渐近方差。接着,将有效影响函数与去偏机器学习(DML)框架结合,提出方法DML-PPCI:若构造估计方程估计量,称为EE-DML-PPCI;若构造目标学习估计量,则称为TMLE-DML-PPCI。两类估计量的渐近方差均达到所推导的效率界。在估计过程中,有效影响函数的估计至关重要。本文中,有效影响函数亦为奈曼正交得分,依赖于瑞斯表示元与回归函数。针对瑞斯表示元估计,我们提出了半监督广义瑞斯回归,并提供收敛率保证。
原文摘要 · Abstract (English)
This study investigates semiparametric efficient estimation of causal and structural parameters in a semi-supervised setting. In our setting, unlabeled auxiliary regressors are available in addition to labeled observations consisting of outcomes and regressors. Our goal is to construct estimators of causal and structural parameters whose asymptotic variances are smaller than those of estimators constructed using only labeled data. We refer to this framework as prediction-powered causal inference (PPCI). We first derive the efficient influence function and the efficiency bound, which imply that the use of auxiliary regressors can attain a smaller asymptotic variance than the efficiency bound attainable from labeled observations alone. Then, by combining the efficient influence function with the debiased machine learning (DML) framework, we propose methods that we call DML-PPCI. If we construct an estimating-equation estimator, we refer to the method as EE-DML-PPCI; if we construct a targeted-learning estimator, we refer to the method as TMLE-DML-PPCI. The asymptotic variances of both estimators match our derived efficiency bound. In the construction of the estimators, estimation of the efficient influence function plays an important role. In our study, the efficient influence function is also a Neyman orthogonal score, which depends on the Riesz representer and the regression function. For Riesz representer estimation, we develop semi-supervised generalized Riesz regression with convergence rate guarantees.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。