通过数据依赖的粗粒化改进IPW估计,提升因果推断鲁棒性。
Smaller Confidence Intervals From IPW Estimators via Data-Dependent Coarsening
- 用共现空间粗化方法构造新型加权估计器,融合现有变体
- 在误差存在时,置信区间大小随ε+1/√n缩放,不再发散
- 适用于高维数据中倾向得分不准确的因果分析场景
逆倾向得分加权(IPW)估计器广泛用于观测研究中的平均处理效应估计。在无混杂假设下,若倾向得分准确且样本数为n,IPW估计器的置信区间大小随n缩小,部分变体可进一步提升缩放速率。然而,现有方法对误差极不鲁棒:即使单个协变量的倾向得分存在ε>0的加性误差,置信区间大小也可能任意增大。此外,即便无误差,在极端倾向得分(接近0或1)存在时,置信区间收敛速度也可能任意缓慢。本文提出一类粗化IPW(CIPW)估计器,涵盖现有各类变体。每个CIPW估计器基于协变量空间的粗化版本进行加权,将某些协变量合并。在适度假设下(如期望结果的Lipschitz连续性和极端倾向得分稀疏性),我们给出高效算法以寻找鲁棒估计器:当倾向得分有ε-不准确且样本数为n时,其置信区间大小缩放为ε+1/√n。相比之下,在相同假设下,现有估计器的置信区间大小始终为Ω(1),与ε和n无关。关键在于,该估计器是数据依赖的;我们证明,任何数据独立的CIPW估计器都无法对误差保持鲁棒。
原文摘要 · Abstract (English)
Inverse propensity-score weighted (IPW) estimators are prevalent in causal inference for estimating average treatment effects in observational studies. Under unconfoundedness, given accurate propensity scores and $n$ samples, the size of confidence intervals of IPW estimators scales down with $n$, and, several of their variants improve the rate of scaling. However, neither IPW estimators nor their variants are robust to inaccuracies: even if a single covariate has an $\varepsilon>0$ additive error in the propensity score, the size of confidence intervals of these estimators can increase arbitrarily. Moreover, even without errors, the rate with which the confidence intervals of these estimators go to zero with $n$ can be arbitrarily slow in the presence of extreme propensity scores (those close to 0 or 1). We introduce a family of Coarse IPW (CIPW) estimators that captures existing IPW estimators and their variants. Each CIPW estimator is an IPW estimator on a coarsened covariate space, where certain covariates are merged. Under mild assumptions, e.g., Lipschitzness in expected outcomes and sparsity of extreme propensity scores, we give an efficient algorithm to find a robust estimator: given $\varepsilon$-inaccurate propensity scores and $n$ samples, its confidence interval size scales with $\varepsilon+1/\sqrt{n}$. In contrast, under the same assumptions, existing estimators' confidence interval sizes are $Ω(1)$ irrespective of $\varepsilon$ and $n$. Crucially, our estimator is data-dependent and we show that no data-independent CIPW estimator can be robust to inaccuracies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。