arXiv:2601.21014stat.MLcs.LG2026-01被引 1

通过模块化子图投票实现高效因果结构学习

Efficient Causal Structure Learning via Modular Subgraph Integration

  • 将全局因果学习分解为基于马尔可夫毯的局部子图
  • 采用加权投票机制提升准确率,显著降低计算开销
  • 支持任意模型与并行计算,适合高维数据场景

从观测数据中学习因果结构仍是基础但计算密集的任务,尤其在高维情形下,现有方法面临搜索空间超指数增长和计算需求上升的问题。为此,我们提出VISTA(基于投票的子图拓扑集成以保证无环性),一种模块化框架,将全局因果结构学习问题基于马尔可夫毯分解为局部子图。通过加权投票机制实现全局整合:利用指数衰减惩罚低支持边,自适应阈值过滤不可靠边,并通过反馈弧集(FAS)算法确保无环性。该框架模型无关,不依赖基学习器的归纳偏置,兼容任意数据设置且无需特定结构形式,完全支持并行化。我们还理论证明了VISTA的有限样本误差界及其在温和条件下的渐近一致性。在合成与真实数据集上的大量实验表明,VISTA在多种基学习器上均显著提升了准确率与效率。

原文摘要 · Abstract (English)

Learning causal structures from observational data remains a fundamental yet computationally intensive task, particularly in high-dimensional settings where existing methods face challenges such as the super-exponential growth of the search space and increasing computational demands. To address this, we introduce VISTA (Voting-based Integration of Subgraph Topologies for Acyclicity), a modular framework that decomposes the global causal structure learning problem into local subgraphs based on Markov Blankets. The global integration is achieved through a weighted voting mechanism that penalizes low-support edges via exponential decay, filters unreliable ones with an adaptive threshold, and ensures acyclicity using a Feedback Arc Set (FAS) algorithm. The framework is model-agnostic, imposing no assumptions on the inductive biases of base learners, is compatible with arbitrary data settings without requiring specific structural forms, and fully supports parallelization. We also theoretically establish finite-sample error bounds for VISTA, and prove its asymptotic consistency under mild conditions. Extensive experiments on both synthetic and real datasets consistently demonstrate the effectiveness of VISTA, yielding notable improvements in both accuracy and efficiency over a wide range of base learners.

因果学习结构学习模块化高效计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。