轻量框架降低因果发现计算开销,无需精确初始结构仍保持高精度。
Causal Discovery for Cross-Sectional Data Based on Super-Structure and Divide-and-Conquer
- 用弱约束结构结合分治策略,减少对条件独立检验的依赖。
- 在多个数据集上比传统方法少用大量条件独立检验,准确率相当。
- 适合缺乏先验知识的大规模生物与社科数据因果分析。
本文针对基于超结构的分治式因果发现中的关键瓶颈——构建精确超结构的高计算成本问题,提出一种轻量级新框架。该框架放宽了对超结构构建的严格要求,同时保留分治算法的优势。通过集成弱约束超结构与高效的图划分和合并策略,显著降低条件独立(CI)检验次数,且不牺牲准确性。我们在具体因果发现算法中实现该框架,并在合成数据上严谨评估各组件性能。在高斯贝叶斯网络(包括magic-NIAB、ECOLI70、magic-IRRI)上的全面实验表明,本方法在结构准确性上可媲美或接近PC与FCI算法,但所需CI测试次数大幅减少。在真实世界中国健康与养老追踪调查(CHARLS)数据集上的验证进一步确认其实际应用价值。结果表明,在对初始超结构假设极少的情况下,仍可实现准确且可扩展的因果发现,为生物医学与社会科学等知识稀缺领域的分治方法应用开辟新路径。
原文摘要 · Abstract (English)
This paper tackles a critical bottleneck in Super-Structure-based divide-and-conquer causal discovery: the high computational cost of constructing accurate Super-Structures--particularly when conditional independence (CI) tests are expensive and domain knowledge is unavailable. We propose a novel, lightweight framework that relaxes the strict requirements on Super-Structure construction while preserving the algorithmic benefits of divide-and-conquer. By integrating weakly constrained Super-Structures with efficient graph partitioning and merging strategies, our approach substantially lowers CI test overhead without sacrificing accuracy. We instantiate the framework in a concrete causal discovery algorithm and rigorously evaluate its components on synthetic data. Comprehensive experiments on Gaussian Bayesian networks, including magic-NIAB, ECOLI70, and magic-IRRI, demonstrate that our method matches or closely approximates the structural accuracy of PC and FCI while drastically reducing the number of CI tests. Further validation on the real-world China Health and Retirement Longitudinal Study (CHARLS) dataset confirms its practical applicability. Our results establish that accurate, scalable causal discovery is achievable even under minimal assumptions about the initial Super-Structure, opening new avenues for applying divide-and-conquer methods to large-scale, knowledge-scarce domains such as biomedical and social science research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。