arXiv:2410.15655cs.LGstat.ME2024-10

解决目标人群新变量缺失下的因果效应估计问题,提升估计精度。

Accounting for Missing Covariates in Heterogeneous Treatment Estimation

  • 基于生态推断思想,利用共现变量约束构建条件效应边界。
  • 提出偏差校正估计器,实现快速收敛与渐近正态性保证。
  • 适用于存在异质处理效应且部分协变量缺失的现实场景。

许多因果推断应用需要将研究人群的处理效应估计结果推广到独立的目标人群。本文考虑一个挑战性场景:目标人群中存在研究中未观测到的协变量。目标是估计在这些新观测协变量条件下,异质处理效应的最紧可能边界。我们提出一种新颖的部分识别策略,灵感来自生态推断;核心思想是:对完整协变量集的条件处理效应估计,在仅限于两人群共同观测协变量时,必须正确地边缘化。此外,我们引入了一种偏差校正的边界估计器,并证明其具有快速收敛速率和统计保障(如渐近正态性)。在真实与合成数据上的实验表明,该框架可产生远比传统方法更紧的边界。

原文摘要 · Abstract (English)

Many applications of causal inference require using treatment effects estimated on a study population to make decisions in a separate target population. We consider the challenging setting where there are covariates that are observed in the target population that were not seen in the original study. Our goal is to estimate the tightest possible bounds on heterogeneous treatment effects conditioned on such newly observed covariates. We introduce a novel partial identification strategy based on ideas from ecological inference; the main idea is that estimates of conditional treatment effects for the full covariate set must marginalize correctly when restricted to only the covariates observed in both populations. Furthermore, we introduce a bias-corrected estimator for these bounds and prove that it enjoys fast convergence rates and statistical guarantees (e.g., asymptotic normality). Experimental results on both real and synthetic data demonstrate that our framework can produce bounds that are much tighter than would otherwise be possible.

因果推断异质效应边界估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。