提出新框架,让因果推断在假设不成立时仍能判断效应方向与稳健性。
Data Fusion for Partial Identification of Causal Effects
- 引入可解释的敏感性参数,量化假设违背程度
- 给出因果效应边界,证明项目STAR结果在假设失效下仍稳健
- 适合关注真实世界因果推断可靠性的研究者
数据融合技术通过整合异质数据源信息,提升学习、泛化与决策能力。在因果推断中,这些方法利用丰富观测数据改善因果效应估计,同时保持对随机对照试验的信任。现有方法常通过假设反事实结果在数据源间可交换来放宽无未观测混杂的强假设,但当两类假设同时失效(实践中常见)时,现有方法无法识别或估计因果效应。本文提出一种新型部分识别框架,使研究者能回答关键问题:因果效应是正向还是负向?假设违背需多严重才会推翻结论?该方法引入可解释的敏感性参数以量化假设违背,并推导对应的因果效应边界。我们开发了双重鲁棒估计器来估计这些边界,并采用崩溃前沿分析法揭示因果结论随假设违背程度变化的情况。将该框架应用于研究班级规模对三年级标准化考试成绩影响的Project STAR研究,结果显示,即使关键假设同时失效,其结果在整体及多个子群体上依然稳健,增强了研究结论的可信度,尽管数据中可能存在未测量偏差。
原文摘要 · Abstract (English)
Data fusion techniques integrate information from heterogeneous data sources to improve learning, generalization, and decision making across data sciences. In causal inference, these methods leverage rich observational data to improve causal effect estimation, while maintaining the trustworthiness of randomized controlled trials. Existing approaches often relax the strong no unobserved confounding assumption by instead assuming exchangeability of counterfactual outcomes across data sources. However, when both assumptions simultaneously fail - a common scenario in practice - current methods cannot identify or estimate causal effects. We address this limitation by proposing a novel partial identification framework that enables researchers to answer key questions such as: Is the causal effect positive or negative? and How severe must assumption violations be to overturn this conclusion? Our approach introduces interpretable sensitivity parameters that quantify assumption violations and derives corresponding causal effect bounds. We develop doubly robust estimators for these bounds and operationalize breakdown frontier analysis to understand how causal conclusions change as assumption violations increase. We apply our framework to the Project STAR study, which investigates the effect of classroom size on students' third-grade standardized test performance. Our analysis reveals that the Project STAR results are robust to simultaneous violations of key assumptions, both on average and across various subgroups of interest. This strengthens confidence in the study's conclusions despite potential unmeasured biases in the data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。