arXiv:2502.06335cs.LG2025-02ICLR被引 9

提出新型采样方法,让贝叶斯神经网络推理更快更稳定。

Microcanonical Langevin Ensembles: Advancing the Sampling of Bayesian Neural Networks

  • 结合优化策略与微正则朗之万采样,构建集成采样框架。
  • 相比最先进方法提速最高达10倍,预测精度与不确定性量化不降反升。
  • 采样耗时可预测,适合并行计算,适用于多种任务和数据类型。

尽管已有进展,贝叶斯神经网络(BNN)的基于采样的推断在概率深度学习中仍是重大挑战。尽管采样方法无需变分分布假设,但现有最先进采样器仍难以有效处理BNN复杂且高度多模态的后验分布。结果是,即使对小型神经网络,采样推理时间仍远长于非贝叶斯方法,即便软件实现已趋高效。除难找到高概率区域外,采样器达到充分探索所需时间也难以预测。为此,我们引入一种集成方法,融合优化策略与近期提出的微正则朗之万蒙特卡洛(MCLMC)采样器,实现高效、鲁棒且可预测的采样性能。相比最先进的无回头采样器(No-U-Turn Sampler),本方法在多种任务和数据模态下实现了最高达一个数量级的提速,同时保持或提升了预测性能与不确定性量化能力。所提微正则朗之万集成及对MCLMC的改进进一步增强了资源需求的可预测性,便于并行化。总体而言,该方法为BNN的实用化、可扩展推断提供了有前景的方向。

原文摘要 · Abstract (English)

Despite recent advances, sampling-based inference for Bayesian Neural Networks (BNNs) remains a significant challenge in probabilistic deep learning. While sampling-based approaches do not require a variational distribution assumption, current state-of-the-art samplers still struggle to navigate the complex and highly multimodal posteriors of BNNs. As a consequence, sampling still requires considerably longer inference times than non-Bayesian methods even for small neural networks, despite recent advances in making software implementations more efficient. Besides the difficulty of finding high-probability regions, the time until samplers provide sufficient exploration of these areas remains unpredictable. To tackle these challenges, we introduce an ensembling approach that leverages strategies from optimization and a recently proposed sampler called Microcanonical Langevin Monte Carlo (MCLMC) for efficient, robust and predictable sampling performance. Compared to approaches based on the state-of-the-art No-U-Turn Sampler, our approach delivers substantial speedups up to an order of magnitude, while maintaining or improving predictive performance and uncertainty quantification across diverse tasks and data modalities. The suggested Microcanonical Langevin Ensembles and modifications to MCLMC additionally enhance the method's predictability in resource requirements, facilitating easier parallelization. All in all, the proposed method offers a promising direction for practical, scalable inference for BNNs.

贝叶斯神经网络采样方法不确定性量化高效推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。