研究深度神经网络在大规模数据下的贝叶斯推断,发现深度影响模型证据的条件。
Bayesian Inference with Shaped Deep Non-linear MLPs
- 基于神经协方差随机微分方程,分析多层感知机在大尺寸下的贝叶斯后验。
- 当 $LP/N$ 为常数时,深度提升可显著增强模型证据,尤其对特定数据生成过程。
- 首次证明贝叶斯后验等价于数据相关核方法,简化了复杂模型的推断分析。
深度学习理论的核心目标之一是刻画神经网络在模型规模和训练集大小同时趋于无穷时的预测行为。由于模型参数量与数据量极限不交换,其极限形式尚不明确。本文研究在训练样本数 $P$、输入维度 $N_0$、隐藏层宽度 $N$ 及隐藏层数 $L$ 均较大的情况下,深度非线性 MLP 的贝叶斯推断。基于 Li 等(2022)提出的神经协方差随机微分方程,我们分析 $LP/N \in \Theta(1)$ 的情形,该比值扮演有效网络深度的角色。框架涵盖光滑与 ReLU 激活函数,并适用于任意温度。我们发现,在 $LP/N$ 的一阶近似下,存在一个简单准则:某些数据生成过程受益于深度,即更大的 $LP/N$ 会提高贝叶斯模型证据。此外,我们给出了物理文献中一个经典结果的全新推导:在 $LP/N$ 的一阶近似下,贝叶斯预测后验极为简洁,等价于数据依赖的核方法。
原文摘要 · Abstract (English)
A central aim of deep learning theory is to characterize how neural networks make predictions in the regime of simultaneously large model and training set size. Since the limits of diverging number of model parameters and dataset size do not commute it is not clear a priori what limits exist. In this work, we shed new light on these questions by studying Bayesian inference in deep non-linear MLPs in the regime where the number of training samples ($P$), the input dimension ($N_0$), the hidden layer width ($N$), and the number of hidden layers ($L$) can all be large. We build on the Neural Covariance SDE (Li et al., 2022) to analyze predictive posteriors in the regime where $LP/N\inΘ(1)$, playing the role of an effective network depth. Our framework covers both smooth and ReLU activation functions and applies to arbitrary temperature. We find to first order in $LP/N$ a simple criterion for which data generating processes benefit from depth in the sense that larger $LP/N$ increases the Bayesian model evidence. We also give a novel derivation of a prior result from the physics literature that at least to first order in $LP/N$, the Bayesian predictive posterior is remarkably simple and is simply equivalent to that of a data-dependent kernel method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。