解决干预实验中的选择偏差问题,提升因果发现准确性。
When Selection Meets Intervention: Additional Complexities in Causal Discovery
- 构建新图模型,同时刻画观测世界与反事实选择机制。
- 在软干预和目标未知数据中,准确识别因果关系与选择机制。
- 适用于药物试验、A/B测试等存在选择偏差的场景。
我们关注干预研究中常被忽视的选择偏差问题:受试者通常仅从特定群体中选取,如药物试验只纳入患者,移动应用的A/B测试仅针对现有用户,基因扰动研究聚焦特定细胞类型(如癌细胞)。忽略此偏差会导致错误的因果发现结果。尽管已有研究意识到该问题,但当前干预因果发现范式仍无法有效应对,因干预发生的时间与地点差异可能引发显著不同的统计模式。为此,我们提出一个图模型,显式建模观测世界(干预实施)与反事实世界(选择发生但干预未施加)的双重结构。我们刻画了该模型的马尔可夫性质,并提出一个可证明正确的算法,从具有软干预和未知目标的数据中,识别因果关系及选择机制至等价类。通过合成数据与真实世界实验验证,本方法在存在选择偏差时仍能有效识别真实因果关系。
原文摘要 · Abstract (English)
We address the common yet often-overlooked selection bias in interventional studies, where subjects are selectively enrolled into experiments. For instance, participants in a drug trial are usually patients of the relevant disease; A/B tests on mobile applications target existing users only, and gene perturbation studies typically focus on specific cell types, such as cancer cells. Ignoring this bias leads to incorrect causal discovery results. Even when recognized, the existing paradigm for interventional causal discovery still fails to address it. This is because subtle differences in when and where interventions happen can lead to significantly different statistical patterns. We capture this dynamic by introducing a graphical model that explicitly accounts for both the observed world (where interventions are applied) and the counterfactual world (where selection occurs while interventions have not been applied). We characterize the Markov property of the model, and propose a provably sound algorithm to identify causal relations as well as selection mechanisms up to the equivalence class, from data with soft interventions and unknown targets. Through synthetic and real-world experiments, we demonstrate that our algorithm effectively identifies true causal relations despite the presence of selection bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。