用集成学习提升大规模贝叶斯网络结构学习的准确性与稳定性。
Scalable Structure Learning of Bayesian Networks by Learning Algorithm Ensembles
- 通过集成多个结构学习算法,构建稳定高效的组合方法。
- 在10,000变量数据集上,准确率提升30%至225%。
- 自动学习最优集成方案,适用于更大规模和不同类型的网络。
从数据中学习贝叶斯网络(BN)结构是一项挑战,尤其在变量数量庞大的情况下。近期提出的分治(D&D)策略为学习大型BN提供了有前景的方法,但仍存在子问题间学习精度不稳定的难题。本文提出结构学习集成(SLE),通过融合多个BN结构学习算法,持续实现高精度。进一步提出自动SLE(Auto-SLE)方法,自动学习近优的SLE,解决手动设计高质量SLE的难题。所学SLE被集成到D&D框架中。大量实验表明,相比使用单一算法的D&D方法,本方法在学习大规模BN时显著更优,在含10,000变量的数据集上准确率普遍提升30%~225%。此外,该方法对包含更多变量(如30,000)及不同网络特性的数据也表现出良好泛化能力。结果表明,自动学习的SLE在可扩展的BN结构学习中具有巨大潜力。
原文摘要 · Abstract (English)
Learning the structure of Bayesian networks (BNs) from data is challenging, especially for datasets involving a large number of variables. The recently proposed divide-and-conquer (D\&D) strategies present a promising approach for learning large BNs. However, they still face a main issue of unstable learning accuracy across subproblems. In this work, we introduce the idea of employing structure learning ensemble (SLE), which combines multiple BN structure learning algorithms, to consistently achieve high learning accuracy. We further propose an automatic approach called Auto-SLE for learning near-optimal SLEs, addressing the challenge of manually designing high-quality SLEs. The learned SLE is then integrated into a D\&D method. Extensive experiments firmly show the superiority of our method over D\&D methods with single BN structure learning algorithm in learning large BNs, achieving accuracy improvement usually by 30\%$\sim$225\% on datasets involving 10,000 variables. Furthermore, our method generalizes well to datasets with many more (e.g., 30000) variables and different network characteristics than those present in the training data for learning the SLE. These results indicate the significant potential of employing (automatic learning of) SLEs for scalable BN structure learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。