提出E-CIT框架,用分治聚合降低因果发现中条件独立检验的计算开销。
Efficient Ensemble Conditional Independence Test Framework for Causal Discovery
- 将数据分块并行执行基础检验,再用稳定分布原理合并p值
- 计算复杂度从高阶降至线性,样本量大时提速显著
- 适用于真实数据复杂场景,适合大规模因果推断研究者
基于约束的因果发现依赖大量条件独立检验(CIT),但其实际应用受限于高昂的计算成本,尤其当检验本身随样本量增长而呈现高时间复杂度时。为解决这一关键瓶颈,本文提出通用且即插即用的集成条件独立检验框架(E-CIT)。E-CIT采用直观的分治-聚合策略:将数据划分为子集,对每个子集独立应用给定的基底CIT,再通过一种基于稳定分布性质的新方法聚合得到的p值。该框架在子集大小固定时,将基底CIT的计算复杂度降至与样本量呈线性关系。此外,我们设计的p值合并方法在较弱条件下具备理论一致性保证。实验表明,E-CIT不仅显著降低CIT及因果发现的计算负担,且表现具有竞争力,尤其在复杂测试场景和真实数据集上展现出改进效果。
原文摘要 · Abstract (English)
Constraint-based causal discovery relies on numerous conditional independence tests (CITs), but its practical applicability is severely constrained by the prohibitive computational cost, especially as CITs themselves have high time complexity with respect to the sample size. To address this key bottleneck, we introduce the Ensemble Conditional Independence Test (E-CIT), a general-purpose and plug-and-play framework. E-CIT operates on an intuitive divide-and-aggregate strategy: it partitions the data into subsets, applies a given base CIT independently to each subset, and aggregates the resulting p-values using a novel method grounded in the properties of stable distributions. This framework reduces the computational complexity of a base CIT to linear in the sample size when the subset size is fixed. Moreover, our tailored p-value combination method offers theoretical consistency guarantees under mild conditions on the subtests. Experimental results demonstrate that E-CIT not only significantly reduces the computational burden of CITs and causal discovery but also achieves competitive performance. Notably, it exhibits an improvement in complex testing scenarios, particularly on real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。