无需额外假设,直接从观测数据推导出精确的因果效应边界。
Information-Theoretic Causal Bounds under Unmeasured Confounding
- 基于信息论构建数据驱动的偏差上界,仅依赖倾向得分。
- 在多种生成机制下,实证显示边界紧致且有效。
- 适合处理未测量混杂因素的因果推断,尤其适用于复杂数据场景。
我们提出一种数据驱动的信息论框架,用于在未测量混杂条件下对因果效应进行严格的部分识别。现有方法通常依赖于受限假设(如结果有界或离散)、外部输入(如工具变量、代理变量或用户指定的敏感性参数)、完整的结构因果模型设定,或仅关注总体平均效应而忽略协变量条件下的因果效应。我们通过建立新颖的信息论数据驱动差异上界,同时克服了上述四类局限。核心理论贡献表明:观测分布 P(Y | A = a, X = x) 与干预分布 P(Y | do(A = a), X = x) 之间的 f-散度,被仅由倾向得分决定的函数所上界。这一结果使得无需外部敏感性参数、辅助变量、完整结构设定或结果有界假设,即可直接从观测数据中实现条件因果效应的严格部分识别。为实际应用,我们开发了一种满足 Neyman 正交性的半参数估计器(Chernozhukov 等, 2018),即使使用灵活的机器学习方法估计干扰函数,也能保证根号 n 一致推断。模拟研究与真实数据应用(见 GitHub 仓库:https://github.com/yonghanjung/Information-Theretic-Bounds)表明,该框架在广泛的数据生成过程中均能提供紧致且有效的因果边界。
原文摘要 · Abstract (English)
We develop a data-driven information-theoretic framework for sharp partial identification of causal effects under unmeasured confounding. Existing approaches often rely on restrictive assumptions, such as bounded or discrete outcomes; require external inputs (for example, instrumental variables, proxies, or user-specified sensitivity parameters); necessitate full structural causal model specifications; or focus solely on population-level averages while neglecting covariate-conditional effects. We overcome all four limitations simultaneously by establishing novel information-theoretic, data-driven divergence bounds. Our key theoretical contribution shows that the f-divergence between the observational distribution P(Y | A = a, X = x) and the interventional distribution P(Y | do(A = a), X = x) is upper bounded by a function of the propensity score alone. This result enables sharp partial identification of conditional causal effects directly from observational data, without requiring external sensitivity parameters, auxiliary variables, full structural specifications, or outcome boundedness assumptions. For practical implementation, we develop a semiparametric estimator satisfying Neyman orthogonality (Chernozhukov et al., 2018), which ensures root-n consistent inference even when nuisance functions are estimated via flexible machine learning methods. Simulation studies and real-world data applications, implemented in the GitHub repository (https://github.com/yonghanjung/Information-Theretic-Bounds), demonstrate that our framework provides tight and valid causal bounds across a wide range of data-generating processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。