arXiv:2602.08315cs.LGstat.ML2026-02

用流匹配加速因果发现中的条件独立检验,训练一次即可反复使用。

Fast Flow Matching based Conditional Independence Tests for Causal Discovery

  • 基于流匹配构建快速条件独立检验,仅需一次模型训练
  • 在高维条件下仍保持低误报率与高检验效能
  • 适合需要高效因果推断的科研与工业场景

基于约束的因果发现方法依赖大量条件独立(CI)检验,因计算复杂度高而难以实际应用。为此,我们提出流匹配为基础的条件独立检验(FMCIT),利用流匹配的高效性,整个因果发现过程仅需一次模型训练,显著加速推理。数值实验表明,FMCIT在高维调节集下仍能有效控制第一类错误,保持高检验力。进一步将FMCIT集成到两阶段引导式PC骨架学习框架中,称为GPC-FMCIT,结合快速筛选与受控精炼,明确限制CI查询次数,同时保持高统计功效。在合成与真实世界数据上的实验显示,该方法在准确率与效率间取得更优权衡,优于现有CI检验及PC变体。

原文摘要 · Abstract (English)

Constraint-based causal discovery methods require a large number of conditional independence (CI) tests, which severely limits their practical applicability due to high computational complexity. Therefore, it is crucial to design an algorithm that accelerates each individual test. To this end, we propose the Flow Matching-based Conditional Independence Test (FMCIT). The proposed test leverages the high computational efficiency of flow matching and requires the model to be trained only once throughout the entire causal discovery procedure, substantially accelerating causal discovery. According to numerical experiments, FMCIT effectively controls type-I error and maintains high testing power under the alternative hypothesis, even in the presence of high-dimensional conditioning sets. In addition, we further integrate FMCIT into a two-stage guided PC skeleton learning framework, termed GPC-FMCIT, which combines fast screening with guided, budgeted refinement using FMCIT. This design yields explicit bounds on the number of CI queries while maintaining high statistical power. Experiments on synthetic and real-world causal discovery tasks demonstrate favorable accuracy-efficiency trade-offs over existing CI testing methods and PC variants.

因果发现流匹配条件独立高效算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。