用几何方法改进神经网络贝叶斯近似,提升预测可靠性且训练成本更低
Tubular Riemannian Laplace Approximations for Bayesian Neural Networks
- 基于流形几何构建概率管状后验,分离参数方向不确定性
- 在ResNet-18上达到与深度集成相当的校准效果,仅需1/5训练成本
- 适合追求高效高可靠预测的工业级模型部署场景
拉普拉斯近似是神经网络中简单实用的近似贝叶斯推断方法,但其欧氏形式难以处理现代深度模型复杂的各向异性、弯曲损失面及大对称群结构。近期研究提出黎曼与几何高斯近似以适应该结构。本文提出管状黎曼拉普拉斯(TRL)近似,显式将后验建模为由函数对称性诱导的低损失山谷中的概率管状结构,使用Fisher/高斯-牛顿度量分离先验主导的切向不确定性与数据主导的横向不确定性。我们将TRL视为一种可扩展的重参数化高斯近似,利用隐式曲率估计在高维参数空间中运行。在ResNet-18(CIFAR-10和CIFAR-100)上的实证评估表明,TRL实现了优异的校准性能,其ECE指标与深度集成相当或更优,同时仅需约1/5的训练成本。TRL有效弥合了单模型效率与集成级可靠性之间的差距。
原文摘要 · Abstract (English)
Laplace approximations are among the simplest and most practical methods for approximate Bayesian inference in neural networks, yet their Euclidean formulation struggles with the highly anisotropic, curved loss surfaces and large symmetry groups that characterize modern deep models. Recent work has proposed Riemannian and geometric Gaussian approximations to adapt to this structure. Building on these ideas, we introduce the Tubular Riemannian Laplace (TRL) approximation. TRL explicitly models the posterior as a probabilistic tube that follows a low-loss valley induced by functional symmetries, using a Fisher/Gauss-Newton metric to separate prior-dominated tangential uncertainty from data-dominated transverse uncertainty. We interpret TRL as a scalable reparametrised Gaussian approximation that utilizes implicit curvature estimates to operate in high-dimensional parameter spaces. Our empirical evaluation on ResNet-18 (CIFAR-10 and CIFAR-100) demonstrates that TRL achieves excellent calibration, matching or exceeding the reliability of Deep Ensembles (in terms of ECE) while requiring only a fraction (1/5) of the training cost. TRL effectively bridges the gap between single-model efficiency and ensemble-grade reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。