arXiv:2606.13818cs.LG2026-06

用贝叶斯方法解析深度学习的泛化与不确定性问题

Uncertainty Estimation and Generalization Bounds for Modern Deep Learning

  • 提出可扩展的贝叶斯框架DVIP,结合隐式过程与深度网络
  • 通过后处理方法为预训练网络注入校准过的不确定性估计
  • 从概率视角统一解释过参数化网络为何能良好泛化

本论文研究贝叶斯原理如何深化对现代深度学习系统的理解。尽管神经网络在预测性能上表现卓越,但其泛化能力与不确定性量化机制仍不完全明晰。本文从方法与理论双重视角推进:将贝叶斯推断、函数空间建模与大偏差理论统一于共同的概率框架下。方法层面,提出可扩展的深层变分隐式过程(DVIP),将隐式过程拓展至深度架构;同时提出两种后处理方法——变分线性拉普拉斯近似(VaLLA)与固定均值高斯过程(FMGP),为预训练确定性网络赋予校准的不确定性估计。理论方面,聚焦机器学习的核心开放问题:为何大规模过参数化神经网络仍能良好泛化?本文构建统一的概率框架,将多样性、平滑性与随机性三种关键机制纳入PAC-Bayesian与大偏差理论的语言体系中。

原文摘要 · Abstract (English)

This thesis investigates how Bayesian principles can deepen our understanding of modern deep learning systems. While neural networks achieve remarkable predictive performance, their ability to generalize and to quantify uncertainty remains only partly understood. This thesis approaches this challenge from both methodological and theoretical angles: unifying Bayesian inference, function-space modeling, and large-deviation theory under a common probabilistic perspective. On the methodological side, the thesis introduces the Deep Variational Implicit Process (DVIP), a scalable Bayesian framework that extends implicit processes to deep architectures. Complementing this, two post-hoc methods -- the Variational Linearized Laplace Approximation (VaLLA) and the Fixed-Mean Gaussian Process (FMGP) -- are proposed to equip pretrained deterministic networks with calibrated uncertainty estimates. The theoretical contributions focus on one of the central open questions in modern machine learning: why do large, over-parameterized neural networks generalize so well? To address this, the thesis develops a unified probabilistic framework that connects three key mechanisms -- diversity, smoothness, and stochasticity -- within the language of PAC-Bayesian and large-deviation theory.

贝叶斯学习不确定性泛化理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。