用干预数据实现大规模因果图推断,还能量化不确定性。
Large-Scale Bayesian Causal Discovery with Interventional Data
- 基于因果效应矩阵的贝叶斯框架,高效建模干预数据
- 在521基因的Perturb-seq数据上实现更优的因果结构恢复
- 可输出边的后验包含概率,适合需要可信推断的研究
以有向无环图(DAG)形式推断变量间的因果关系是一个重要但极具挑战性的问题。近年来,高通量基因组扰动筛查的发展推动了利用干预数据提升模型识别能力的方法进展。然而,现有方法在大规模任务上表现不佳,且无法量化不确定性。本文提出干预贝叶斯因果发现(IBCD),一种基于干预数据的经验贝叶斯框架。该方法对总因果效应矩阵进行建模,其可通过矩阵正态分布近似,而非直接建模完整数据矩阵。在边结构上采用尖峰-平滑霍舍因子先验,并分别从观测数据中学习无标度和Erdős-Rényi结构的数据驱动权重,将每条边视为潜在变量以实现不确定性感知推断。大量模拟实验表明,IBCD在结构恢复上优于现有基线方法。我们将IBCD应用于521个基因的CRISPR扰动(Perturb-seq)数据,证明边的后验包含概率能有效识别稳健的图结构。
原文摘要 · Abstract (English)
Inferring the causal relationships among a set of variables in the form of a directed acyclic graph (DAG) is an important but notoriously challenging problem. Recently, advancements in high-throughput genomic perturbation screens have inspired development of methods that leverage interventional data to improve model identification. However, existing methods still suffer poor performance on large-scale tasks and fail to quantify uncertainty. Here, we propose Interventional Bayesian Causal Discovery (IBCD), an empirical Bayesian framework for causal discovery with interventional data. Our approach models the likelihood of the matrix of total causal effects, which can be approximated by a matrix normal distribution, rather than the full data matrix. We place a spike-and-slab horseshoe prior on the edges and separately learn data-driven weights for scale-free and Erdős-Rényi structures from observational data, treating each edge as a latent variable to enable uncertainty-aware inference. Through extensive simulation, we show that IBCD achieves superior structure recovery compared to existing baselines. We apply IBCD to CRISPR perturbation (Perturb-seq) data on 521 genes, demonstrating that edge posterior inclusion probabilities enable identification of robust graph structures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。