arXiv:2510.23684stat.MLcs.LG2025-10NeurIPS被引 3

提出新变分推断方法,让深度网络贝叶斯推理更稳定高效

VIKING: Deep variational inference with stochastic projections

论文配图:VIKING: Deep variational inference with stochastic projections
图 1 · 摘自论文原文
  • 通过两个独立线性子空间建模参数空间,分离数据内外的函数变化
  • 在多个模型和数据集上实现当前最优的预测性能与校准效果
  • 方法简洁可扩展,适合希望提升深度网络不确定性估计的研究者

变分均值场近似在现代过参数化深度神经网络中表现不佳。尽管贝叶斯方法理论上应提供高质量预测与不确定性估计,但实际中常出现训练不稳定、预测能力差和校准不足的问题。本文基于神经网络重参数化的最新进展,提出一种简单的变分族,考虑参数空间中两个独立的线性子空间,分别代表训练数据支撑集内和外的功能变化。该方法构建了一个完全相关的近似后验,反映过参数化特性,并能调节易于理解的超参数。我们开发了可扩展的数值算法,用于最大化对应的证据下界(ELBO)并从近似后验中采样。实验表明,在多种任务、模型和数据集上,该方法的表现优于广泛基准方法。结果表明,只要设计合理的推断机制以反映重参数化几何结构,深度神经网络的近似贝叶斯推断远未陷入困境。

原文摘要 · Abstract (English)

Variational mean field approximations tend to struggle with contemporary overparametrized deep neural networks. Where a Bayesian treatment is usually associated with high-quality predictions and uncertainties, the practical reality has been the opposite, with unstable training, poor predictive power, and subpar calibration. Building upon recent work on reparametrizations of neural networks, we propose a simple variational family that considers two independent linear subspaces of the parameter space. These represent functional changes inside and outside the support of training data. This allows us to build a fully-correlated approximate posterior reflecting the overparametrization that tunes easy-to-interpret hyperparameters. We develop scalable numerical routines that maximize the associated evidence lower bound (ELBO) and sample from the approximate posterior. Empirically, we observe state-of-the-art performance across tasks, models, and datasets compared to a wide array of baseline methods. Our results show that approximate Bayesian inference applied to deep neural networks is far from a lost cause when constructing inference mechanisms that reflect the geometry of reparametrizations.

变分推断贝叶斯深度学习过参数化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。