在存在隐变量和选择偏差时,高效准确地发现目标变量的直接因果关系。
Local Causal Structure Learning in the Presence of Latent Variables and Selection Bias

- 聚焦局部区域,无需构建全局因果图
- 在真实数据条件下保持高结构准确率,计算成本远低于全局方法
- 适用于基因表达等大规模生物数据分析
从观测数据中发现目标变量的直接因果关系是因果发现中的基础问题,广泛应用于基因调控分析和生物医学研究。现有方法或需学习全局因果结构,计算开销大;或假设无隐变量和选择偏差,但这些假设在实际中常不成立。针对此问题,本文研究在存在隐变量和选择偏差下的局部因果结构学习。首先刻画了支持目标特定因果发现的局部区域;进而建立该局部区域上观测分布所蕴含的因果信息与全局结构间的理论联系。基于此,提出LoCaLS算法,在标准假设下具有完备性与正确性,可识别出与全局方法相同的直接因果关系,同时允许隐变量和选择偏差存在。大量实验表明,该方法在随机和真实结构上均显著优于现有局部方法,且计算代价远低于主流全局方法。对两个真实基因表达数据集的应用揭示了符合生物学常识的目标特异性因果结构,验证了其在大规模生物数据分析中的实用性。
原文摘要 · Abstract (English)
Discovering the direct causes and effects of a target variable from observational data is a fundamental problem in causal discovery, with broad applications in domains such as gene regulatory analysis and biomedical research. Existing causal discovery methods either learn a global causal structure, which incurs substantial computational cost, or assume the absence of latent variables and selection bias, assumptions that are often violated in real-world settings. Motivated by these challenges, we study local causal structure learning in the presence of latent variables and selection bias. Specifically, we first characterize a local region that enables target-specific causal discovery without recovering the entire global structure. We then establish a theoretical bridge between causal information learned from the observed distribution induced on this local region and the corresponding information in the global causal structure. Building on these foundations, we propose LoCaLS, a local causal structure learning algorithm that is sound and complete under standard assumptions and identifies the same direct causes and effects of a target variable as those identifiable by global causal discovery methods, while allowing for latent variables and selection bias. Extensive experiments on random and real-world structures demonstrate that the proposed method consistently achieves higher structural accuracy than existing local methods while requiring substantially less computational effort than state-of-the-art global methods. Furthermore, applications to two real-world gene expression datasets reveal biologically plausible target-specific causal structures, demonstrating its practical applicability in large-scale biological data analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。