arXiv:2505.11085cs.LGcs.AI2025-05被引 2

提出FastKCI,让大规模数据的因果推断更快更高效

A Fast Kernel-based Conditional Independence test with Application to Causal Discovery

  • 用高斯混合模型分治数据,分块并行做局部KCI检验
  • 在真实数据上比原方法快数倍,统计功效几乎不变
  • 适合处理海量数据的因果发现任务,尤其适合工业级应用

基于核函数的条件独立性(KCI)测试是一种强大的非参数方法,广泛用于因果发现。尽管具有灵活性和统计可靠性,其立方级计算复杂度限制了在大规模数据上的应用。为此,我们提出FastKCI,一种可扩展且可并行的核基条件独立性测试方法。该方法受高斯过程中显式并行推理启发,采用专家混合策略:基于条件变量的高斯混合模型对数据进行分区,在各子集上并行执行局部KCI测试,并通过重要性加权采样聚合结果。在合成数据与真实世界生产数据基准上的实验表明,FastKCI在保持原始KCI测试统计功效的同时,实现了显著的计算加速。因此,FastKCI为大规模数据上的因果推断提供了实用高效的解决方案。

原文摘要 · Abstract (English)

Kernel-based conditional independence (KCI) testing is a powerful nonparametric method commonly employed in causal discovery tasks. Despite its flexibility and statistical reliability, cubic computational complexity limits its application to large datasets. To address this computational bottleneck, we propose \textit{FastKCI}, a scalable and parallelizable kernel-based conditional independence test that utilizes a mixture-of-experts approach inspired by embarrassingly parallel inference techniques for Gaussian processes. By partitioning the dataset based on a Gaussian mixture model over the conditioning variables, FastKCI conducts local KCI tests in parallel, aggregating the results using an importance-weighted sampling scheme. Experiments on synthetic datasets and benchmarks on real-world production data validate that FastKCI maintains the statistical power of the original KCI test while achieving substantial computational speedups. FastKCI thus represents a practical and efficient solution for conditional independence testing in causal inference on large-scale data.

因果发现核方法并行计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。