系统梳理因果发现中条件独立检验方法与挑战
Conditional Independence Tests for Constraint-Based Causal Discovery: A Survey
- 按六类方法归类条件独立检验,分析适用场景与局限
- 揭示测试误差如何导致因果图骨架与方向错误
- 适合从事因果推断、生物医学数据挖掘的研究者
条件独立(CI)检验是基于约束的因果发现算法的核心统计工具,如PC和FCI算法中的图结构剪枝与方向推断均依赖于CI判断。本综述聚焦高维及混合类型数据场景下(常见于生物医学领域)的CI检验,系统整理了六类主流方法:偏相关、列联表、回归、近邻、核方法及基于机器学习的方法。重点分析各方法在不同假设下的稳健性,探讨其在反映数据生成分布时的有效性与失效条件。进一步将测试层面的特性(如条件集增大导致检验效能下降、第一类/第二类错误不对称影响)与图层面的错误(骨架恢复与V型结构方向推断)关联。还对比了主流R与Python库中的方法实现,并总结开放挑战:无需离散化的混合类型CI检验、小样本误差控制、以及提升CI检验可扩展性的策略。
原文摘要 · Abstract (English)
Conditional Independence (CI) tests are the statistical engine of constraint-based causal discovery: in algorithms such as PC (Peter-Clark) and FCI (Fast Causal Inference), skeleton pruning and key orientations follow directly from CI decisions. This survey reviews CI testing with emphasis on assumptions, robustness, and scalability in high-dimensional and mixed-type settings common in biomedical domains. The survey organizes widely used CI methods into six families: partial-correlation, contingency-table, regression, nearest-neighbor, kernel, and machine-learning-based. Special emphasis is provided on the robustness layers that address the limitations of these families. For each family, the survey examines when CI decisions reflect the data-generating distribution and when they fail. By this, we link test-level properties, including power decay with conditioning set size and asymmetric type I/II error consequences, to graph-level errors in skeleton recovery and v-structure orientation. The survey also compares adoption across major R and Python libraries and summarizes open challenges, including mixed-type CI testing without discretization, small-sample error control, and strategies for improving scalability of CI-testing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。