提出新方法解决干预数据中选择偏差对因果推断的干扰问题。
Characterization and Learning of Causal Graphs with Latent Confounders and Post-treatment Selection from Interventional Data
- 显式建模干预后选择机制,区分因果关系与选择模式。
- 在合成与真实数据上成功恢复因果结构,即使存在隐变量和选择偏差。
- 适合生物医学等含复杂筛选流程的研究场景使用。
干预式因果发现通过干预引发的分布变化来识别因果关系,即使存在隐变量。我们指出一个常见但常被忽视的挑战:干预后选择,即干预后样本按特定标准筛选进入数据集。这在基因表达分析等生物研究中普遍存在,例如仅保留高活性细胞。忽略此过程可能导致虚假依赖和分布变化,伪装成因果效应,从而扭曲因果发现结果。为此,我们提出一种新的因果形式化,显式建模干预后选择,并揭示其对干预的差异化反应可区分因果关系与选择模式,突破传统等价类限制,逼近真实因果结构。我们刻画了其马尔可夫性质,提出细粒度干预等价类FI-Markov等价,用新图示F-PAG表示。进一步设计了可证明完全且正确的算法F-FCI,利用观测与干预数据,识别因果关系、隐变量及干预后选择,至$φ$-Markov等价。在合成与真实数据上的实验表明,该方法在存在选择与隐变量时仍能准确恢复因果关系。
原文摘要 · Abstract (English)
Interventional causal discovery seeks to identify causal relations by leveraging distributional changes introduced by interventions, even in the presence of latent confounders. Beyond the spurious dependencies induced by latent confounders, we highlight a common yet often overlooked challenge in the problem due to post-treatment selection, in which samples are selectively included in datasets after interventions. This fundamental challenge widely exists in biological studies; for example, in gene expression analysis, both observational and interventional samples are retained only if they meet quality control criteria (e.g., highly active cells). Neglecting post-treatment selection may introduce spurious dependencies and distributional changes under interventions, which can mimic causal responses, thereby distorting causal discovery results and challenging existing causal formulations. To address this, we introduce a novel causal formulation that explicitly models post-treatment selection and reveals how its differential reactions to interventions can distinguish causal relations from selection patterns, allowing us to go beyond traditional equivalence classes toward the underlying true causal structure. We then characterize its Markov properties and propose a Fine-grained Interventional equivalence class, named FI-Markov equivalence, represented by a new graphical diagram, F-PAG. Finally, we develop a provably sound and complete algorithm, F-FCI, to identify causal relations, latent confounders, and post-treatment selection up to $\mathcal{FI}$-Markov equivalence, using both observational and interventional data. Experimental results on synthetic and real-world datasets demonstrate that our method recovers causal relations despite the presence of both selection and latent confounders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。