提出新方法让大规模因果发现更快更准。
Relaxed Sparsest-Permutation Formulation for Causal Discovery at Scale

- 用松弛的稀疏三角分解替代精确分解,降低计算负担。
- 在真实和合成数据上,准确率接近慢得多的基线方法。
- 适合处理上万变量的大规模因果结构学习任务。
尽管大数据日益普及,大规模因果结构学习仍面临计算瓶颈。本文重新审视线性结构方程模型中的稀疏排列学习,发现结构恢复无需精确的Cholesky分解。由此提出支持级松弛策略,在精度支撑筛选图上搜索稀疏三角因子,并通过掩码零填充不完全Cholesky分解高效评估候选排序。在总体层面,证明了在无抵消和最稀疏马尔可夫表示假设下,该方法能正确恢复马尔可夫等价类(MEC),且对排序误设具有鲁棒性。基于此,我们提出SCOPE——一种基于稀疏Cholesky的流水线,实现该松弛公式。实验表明,SCOPE在合成与真实数据上的MEC恢复准确率媲美显著更慢的基线方法,同时大幅降低运行时间,可扩展至10,000个变量。
原文摘要 · Abstract (English)
Despite the growing availability of large datasets, causal structure learning remains computationally prohibitive at scale. We revisit sparsest-permutation learning for linear structural equation models and show that exact Cholesky factorization is unnecessary for structure recovery. This observation motivates a support-level relaxation that searches for sparse triangular factors over a precision-support screening graph. The relaxed formulation can be efficiently evaluated via masked zero-fill incomplete Cholesky factorization, enabling scalable comparison of candidate orderings. At the population level, we establish soundness for Markov equivalence class (MEC) recovery under no-cancellation and sparsest Markov representation assumptions, as well as robustness to ordering misspecification. Motivated by these guarantees, we introduce SCOPE, a sparse-Cholesky pipeline that provides a scalable implementation of the relaxed formulation. Experiments on synthetic and real datasets demonstrate that SCOPE matches the MEC recovery accuracy of substantially slower baselines, while achieving significantly reduced runtime and scaling to 10k variables.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。