无需因果图即可准确估计处理效应,解决观察数据中的混杂偏倚问题。
Doubly robust identification of treatment effects from multiple environments
- 利用多源数据的异质性,不依赖因果图识别处理效应
- 在治疗或结果变量的父节点可观测时,可实现无偏估计
- 适合医学、社会科学等无法获取完整因果信息的场景
实际和伦理限制常要求在医学和社会科学中使用观察数据进行因果推断,但这类数据易受混杂因素影响,可能破坏因果结论的有效性。尽管若已知潜在因果图可纠正偏差,但在实际中很少可行。常见做法是调整所有可用协变量,但当存在后处理或未观测变量时仍可能导致估计偏误。我们提出RAMEN算法,通过利用多个数据源的异质性,在无需知晓或学习底层因果图的情况下,实现无偏处理效应估计。其关键优势在于双重稳健识别:只要治疗或结果的父节点被观测到,且满足不变性假设,即可识别处理效应。在合成与真实数据集上的实证评估表明,该方法优于现有技术。
原文摘要 · Abstract (English)
Practical and ethical constraints often require the use of observational data for causal inference, particularly in medicine and social sciences. Yet, observational datasets are prone to confounding, potentially compromising the validity of causal conclusions. While it is possible to correct for biases if the underlying causal graph is known, this is rarely a feasible ask in practical scenarios. A common strategy is to adjust for all available covariates, yet this approach can yield biased treatment effect estimates, especially when post-treatment or unobserved variables are present. We propose RAMEN, an algorithm that produces unbiased treatment effect estimates by leveraging the heterogeneity of multiple data sources without the need to know or learn the underlying causal graph. Notably, RAMEN achieves doubly robust identification: it can identify the treatment effect whenever the causal parents of the treatment or those of the outcome are observed, and the node whose parents are observed satisfies an invariance assumption. Empirical evaluations on synthetic and real-world datasets show that our approach outperforms existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。