arXiv:2410.19780stat.MLcs.LG2024-10被引 15

提出新采样算法,提升贝叶斯神经网络后验分布的精准度。

Sampling from Bayesian Neural Network Posteriors with Symmetric Minibatch Splitting Langevin Dynamics

  • 通过对称小批量分裂与朗之万动力学结合,每步仅用一个子批
  • 采样偏差仅随步长平方和维度平方根增长,控制极佳
  • 在多个图像数据集上显著改善模型预测校准性能

我们提出一种可扩展的动能朗之万动力学算法,用于大样本和人工智能应用中的参数空间采样。该方法将对小批量的对称前向/后向遍历与朗之万动力学的对称离散化相结合。针对特定的朗之万分裂方法(UBU),我们证明所得对称小批量分裂-UBU(SMS-UBU)积分器的偏差为 $O(h^2 d^{1/2})$,其中维度 $d>0$,步长 $h>0$,尽管每迭代仅使用一个子批,仍能有效控制采样偏差随步长的变化。我们将该算法应用于探索贝叶斯神经网络(BNNs)后验分布的局部模式,并评估了卷积神经网络架构在三个不同数据集(Fashion-MNIST、Celeb-A 和胸部X光片)上的分类问题中,后验预测概率的校准性能。结果表明,使用SMS-UBU采样的BNN在预测校准方面显著优于标准训练方法和随机权重平均法。

原文摘要 · Abstract (English)

We propose a scalable kinetic Langevin dynamics algorithm for sampling parameter spaces of big data and AI applications. Our scheme combines a symmetric forward/backward sweep over minibatches with a symmetric discretization of Langevin dynamics. For a particular Langevin splitting method (UBU), we show that the resulting Symmetric Minibatch Splitting-UBU (SMS-UBU) integrator has bias $O(h^2 d^{1/2})$ in dimension $d>0$ with stepsize $h>0$, despite only using one minibatch per iteration, thus providing excellent control of the sampling bias as a function of the stepsize. We apply the algorithm to explore local modes of the posterior distribution of Bayesian neural networks (BNNs) and evaluate the calibration performance of the posterior predictive probabilities for neural networks with convolutional neural network architectures for classification problems on three different datasets (Fashion-MNIST, Celeb-A and chest X-ray). Our results indicate that BNNs sampled with SMS-UBU can offer significantly better calibration performance compared to standard methods of training and stochastic weight averaging.

贝叶斯神经网络采样算法模型校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。