arXiv:2607.25312cs.LG2026-07中稿 · presentation at th…

利用优化器自带的曲率信息,加速贝叶斯神经网络采样并提升稳定性。

Guiding Posterior Exploration with Optimizer-Derived Geometry

  • 用AdamW等自适应优化器的曲率信息指导采样方向。
  • 无需额外计算成本,显著缩短采样预热时间。
  • 适用于各类数据集和模型架构,提升不确定性量化效果。

基于采样的方法为贝叶斯神经网络中的不确定性量化提供了理论基础,但其实际应用常受高维多峰后验分布探索的计算成本制约。为克服这一挑战,贝叶斯深度集成(Bayesian Deep Ensembles)通过从多个优化解中热启动采样已被证明是有效策略。本文表明,在AdamW等自适应优化器的热启动过程中,作为副产物计算的曲率估计可低成本地用于指导采样阶段。我们提出的基于优化器衍生几何的预条件采样策略,能显著减少甚至消除长周期的采样预热阶段,并带来更高的数值稳定性。该方法在不增加额外计算成本的前提下,始终维持或提升预测性能与不确定性量化效果。我们在多种数据集和网络架构上验证了结果的一致性。

原文摘要 · Abstract (English)

Sampling-based methods offer a principled approach to uncertainty quantification in Bayesian neural networks. Their practical use, however, is often challenged by the computational cost of exploring high-dimensional and multimodal posterior distributions. To overcome these difficulties, Bayesian Deep Ensembles, i.e., warmstarting the sampling from several optimized solutions, have proven to be an effective strategy. In this paper, we demonstrate that curvature estimates computed during the warmstart as a byproduct in adaptive optimizers such as AdamW can inform the sampling phase at negligible additional cost. Specifically, our proposed preconditioned sampling strategy based on optimizer-derived geometries can substantially reduce or even eliminate the need for a lengthy sampling burn-in phase and leads to greater numerical stability. This approach consistently maintains or improves predictive performance and uncertainty quantification without any additional computational costs. We confirm the consistency of our findings across various datasets and network architectures.

贝叶斯神经网络采样优化曲率信息高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。