arXiv:2506.04194math.STcs.LG2025-06中稿 · presentation at th…被引 3

提出新识别条件,让因果推断在无混淆和重叠不成立时仍可计算平均处理效应。

What Makes Treatment Effects Identifiable? Characterizations and Estimators Beyond Unconfoundedness

  • 基于学习理论,给出ATE可识别的充要条件
  • 在经典回归不连续设计等场景下证明可识别性
  • 为复杂观测研究中的因果推断提供新方法

大多数常用的平均处理效应(ATE)估计器依赖于无混淆性和重叠性假设。无混淆性要求可观测协变量能解释结果与处理之间的所有相关性;重叠性要求所有个体都有随机的处理决策概率。然而,许多研究(如处理决策确定性的回归不连续设计)常违反这些假设。本文首次研究了超越无混淆性和重叠性的一般识别条件,借鉴统计学习理论,提出一个可解释的充要条件,既适用于ATE,也可推广至处理组平均效应(ATT)及其他处理效应。我们展示了在若干经典模型中(如Tan, 2006;Rosenbaum, 2002;Thistlethwaite and Campbell, 1960的回归不连续设计),该条件在温和分布假设下成立,从而证明了这些情形下ATE可识别。同时,在自然附加假设下,我们还证明了可通过有限样本估计ATE。研究成果为学习理论与因果推断的融合开辟了新路径,尤其适用于具有复杂处理机制的观测研究。

原文摘要 · Abstract (English)

Most of the widely used estimators of the average treatment effect (ATE) in causal inference rely on the assumptions of unconfoundedness and overlap. Unconfoundedness requires that the observed covariates account for all correlations between the outcome and treatment. Overlap requires the existence of randomness in treatment decisions for all individuals. Nevertheless, many types of studies frequently violate unconfoundedness or overlap, for instance, observational studies with deterministic treatment decisions - popularly known as Regression Discontinuity designs - violate overlap. In this paper, we initiate the study of general conditions that enable the identification of the average treatment effect, extending beyond unconfoundedness and overlap. In particular, following the paradigm of statistical learning theory, we provide an interpretable condition that is sufficient and necessary for the identification of ATE. Moreover, this condition also characterizes the identification of the average treatment effect on the treated (ATT) and can be used to characterize other treatment effects as well. To illustrate the utility of our condition, we present several well-studied scenarios where our condition is satisfied and, hence, we prove that ATE can be identified in regimes that prior works could not capture. For example, under mild assumptions on the data distributions, this holds for the models proposed by Tan (2006) and Rosenbaum (2002), and the Regression Discontinuity design model introduced by Thistlethwaite and Campbell (1960). For each of these scenarios, we also show that, under natural additional assumptions, ATE can be estimated from finite samples. We believe these findings open new avenues for bridging learning-theoretic insights and causal inference methodologies, particularly in observational studies with complex treatment mechanisms.

因果推断处理效应识别条件回归不连续

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。