arXiv:2602.05873cs.LG2026-02

提出一种可扩展的基于得分的变分推断方法,解决贝叶斯神经网络训练中的模式崩溃问题。

Large-scale Score-based Variational Posterior Inference for Bayesian Deep Neural Networks

  • 采用得分匹配损失与近端惩罚项结合的迭代优化策略
  • 在视觉识别与时间序列预测任务中实现稳定性能,支持大规模网络
  • 无需重参数化采样,支持随机梯度的无偏小批量得分计算

贝叶斯(深度)神经网络在不确定性量化、抗噪声、抗过拟合等方面优于传统点估计深度学习。变分推断(VI)是主流近似推断方法之一。尽管基于ELBO的变分自由能方法占主导地位,本文提出一种基于得分的替代方案,可缓解ELBO方法中常见的模式崩溃问题。虽然已有若干得分型VI方法,但大多不适用于大规模贝叶斯神经网络,受限于计算与技术因素。本文提出一种新型可扩展的变分推断方法:通过在迭代中联合使用得分匹配损失与近端惩罚项,避免重参数化采样,并支持通过随机梯度获得含噪声的无偏小批量得分。这使得方法可扩展至大规模网络,包括视觉变换器(Vision Transformers)。在多个基准任务(如视觉识别与时间序列预测)上,使用大规模深度网络验证了该方法的有效性。

原文摘要 · Abstract (English)

Bayesian (deep) neural networks (BNN) are often more attractive than the vanilla point-estimate deep learning in various aspects including uncertainty quantification, robustness to noise, resistance to overfitting, and more. The variational inference (VI) is one of the most widely adopted approximate inference methods. Whereas the ELBO-based variational free energy method is a dominant choice in the literature, in this paper we introduce a score-based alternative for BNN variational inference. Score-based VI can address the known issue of mode collapsing in ELBO-based VI. Although several score-based VI methods have been proposed in the community, most are not adequate for large-scale BNNs for various computational and technical reasons. We propose a novel scalable VI method where the learning objective combines the score matching loss and the proximal penalty term in iterations, which helps our method avoid the reparametrized sampling, and allows for noisy unbiased mini-batch scores through stochastic gradients. This in turn makes our method scalable to large-scale neural networks including Vision Transformers. On several benchmarks including visual recognition and time-series forecasting with large-scale deep networks, we empirically show the effectiveness of our approach.

贝叶斯神经网络变分推断得分匹配大规模模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。