用信息论设计新模型,高效准确找变量间关键依赖关系。
Neural Autoregressive Flows for Markov Boundary Learning
- 用条件熵做评分,结合掩码自回归网络捕捉复杂依赖
- 多项式时间并行贪心搜索,理论支持且速度快
- 适合需要可靠因果发现的科研与工业场景
恢复马尔可夫边界——即对响应变量预测性能最优的最小变量集——在诸多应用中至关重要。尽管近期方法通过评分局部因果结构改进了传统约束法,但仍依赖非参数估计器和启发式搜索,缺乏可靠性理论保证。本文提出一种高效马尔可夫边界发现框架,将信息论中的条件熵作为评分标准。设计新型掩码自回归网络以捕捉复杂依赖关系,提出可在多项式时间内并行执行的贪心搜索策略,并提供理论分析支持。此外,讨论了用学习到的马尔可夫边界初始化图结构,可加速因果发现收敛。在真实世界与合成数据集上的全面评估表明,该方法在马尔可夫边界发现和因果发现任务中均具有优越性能与可扩展性。
原文摘要 · Abstract (English)
Recovering Markov boundary -- the minimal set of variables that maximizes predictive performance for a response variable -- is crucial in many applications. While recent advances improve upon traditional constraint-based techniques by scoring local causal structures, they still rely on nonparametric estimators and heuristic searches, lacking theoretical guarantees for reliability. This paper investigates a framework for efficient Markov boundary discovery by integrating conditional entropy from information theory as a scoring criterion. We design a novel masked autoregressive network to capture complex dependencies. A parallelizable greedy search strategy in polynomial time is proposed, supported by analytical evidence. We also discuss how initializing a graph with learned Markov boundaries accelerates the convergence of causal discovery. Comprehensive evaluations on real-world and synthetic datasets demonstrate the scalability and superior performance of our method in both Markov boundary discovery and causal discovery tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。