arXiv:2605.10533cs.LG2026-05

提出新方法识别哪些变量是混杂因子,让因果推断更清晰。

ConfoundingSHAP: Quantifying confounding strength in causal inference

论文配图:ConfoundingSHAP: Quantifying confounding strength in causal inference
图 1 · 摘自论文原文
  • 基于谢帕利值设计专属游戏,量化每个变量的混杂强度。
  • 在多个数据集上验证,能准确识别驱动混杂的关键变量。
  • 适合做因果推断的研究者和需要解释性分析的实践者。

在因果推断中,混杂因子是同时影响处理决策和结果的变量。然而,在观察性研究中,处理分配机制未知,难以确定哪些协变量构成混杂因子。本文提出 ConfoundingSHAP,一种基于谢帕利值的方法,用于为个体协变量分配混杂强度。贡献有二:第一,构建针对混杂强度推断的谢帕利博弈,其值不同于标准 SHAP 在因果目标(如治疗效应异质性)上的应用,后者不适用于此任务;第二,由于需评估大量调整集的值函数,我们采用可扩展的 TabPFN 估计方法,避免全量重拟合。在多个数据集上实证表明,ConfoundingSHAP 能有效揭示哪些协变量主导混杂,从而为实际因果推断提供更深入洞见。

原文摘要 · Abstract (English)

In causal inference, confounders are variables that influence both treatment decisions and outcomes. However, unlike as in randomized clinical trials, the treatment assignment mechanism in observational studies is not known, and it is thus unclear which covariates act as confounders. Here, we aim to generate insight for causal inference and answer: which of the observed covariates act as confounders? We introduce ConfoundingSHAP, a Shapley-based method for attributing confounding strength to individual covariates. Our contributions are twofold. First, we propose a Shapley game targeted to infer the confounding strength of the covariates. Our resulting Shapley values differ from the standard applications of SHAP explanations on causal targets, such as understanding treatment effect heterogeneity, which are ill-suited for our task. Second, as our task requires evaluating the value function over many adjustment sets, we provide a scalable TabPFN-based estimation that avoids exhaustive refitting. We demonstrate the practical value across various datasets, where ConfoundingSHAP provides informative explanations of which observed covariates drive confounding and thereby helps to provide more insight for causal inference in practice.

因果推断混杂因子可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。