将因果推断重构为分布适应问题,提升处理效应估计的稳健性。
Causal Inference as Distribution Adaptation: Optimizing ATE Risk under Propensity Uncertainty
- 把平均处理效应估计看作协变量分布偏移下的域自适应问题
- 新方法在模型误设时使均方误差降低15%
- 适合关注因果推断鲁棒性的研究者使用
标准因果推断方法如结果回归和逆概率加权回归调整(IPWRA),通常基于缺失数据插补和识别理论推导。本文从机器学习视角统一这些方法,将平均处理效应(ATE)估计重新建模为存在分布偏移的域自适应问题。我们证明经典Hajek估计量是受限于常数假设类的IPWRA特例,而IPWRA本质上是用于纠正处理组与目标人群间协变量偏移的重要加权经验风险最小化。基于这一统一框架,我们分析双重稳健估计器的优化目标,指出传统方法要求结果模型独立无偏,这是充分但非必要条件。我们定义真正的“ATE风险函数”,并证明只需处理组与对照组模型偏差结构上抵消即可。据此提出联合稳健估计器(JRE):不再分步估计倾向得分与结果模型,而是利用倾向得分的自助法不确定性量化,联合训练结果模型。通过在倾向得分分布上优化期望ATE风险,JRE利用模型自由度增强对倾向得分误设的鲁棒性。模拟研究表明,在有限样本且结果模型误设条件下,JRE相比标准IPWRA可实现最高15%的均方误差降低。
原文摘要 · Abstract (English)
Standard approaches to causal inference, such as Outcome Regression and Inverse Probability Weighted Regression Adjustment (IPWRA), are typically derived through the lens of missing data imputation and identification theory. In this work, we unify these methods from a Machine Learning perspective, reframing ATE estimation as a \textit{domain adaptation problem under distribution shift}. We demonstrate that the canonical Hajek estimator is a special case of IPWRA restricted to a constant hypothesis class, and that IPWRA itself is fundamentally Importance-Weighted Empirical Risk Minimization designed to correct for the covariate shift between the treated sub-population and the target population. Leveraging this unified framework, we critically examine the optimization objectives of Doubly Robust estimators. We argue that standard methods enforce \textit{sufficient but not necessary} conditions for consistency by requiring outcome models to be individually unbiased. We define the true "ATE Risk Function" and show that minimizing it requires only that the biases of the treated and control models structurally cancel out. Exploiting this insight, we propose the \textbf{Joint Robust Estimator (JRE)}. Instead of treating propensity estimation and outcome modeling as independent stages, JRE utilizes bootstrap-based uncertainty quantification of the propensity score to train outcome models jointly. By optimizing for the expected ATE risk over the distribution of propensity scores, JRE leverages model degrees of freedom to achieve robustness against propensity misspecification. Simulation studies demonstrate that JRE achieves up to a 15\% reduction in MSE compared to standard IPWRA in finite-sample regimes with misspecified outcome models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。