突破传统高斯过程极限,揭示贝叶斯神经网络的复杂特征学习机制。
Beyond NNGP: Large Deviations and Feature Learning in Bayesian Neural Networks
- 基于大偏差理论构建变分目标,从函数层面定义模型复杂度。
- 后验输出率函数通过预测器与内部核函数联合优化获得,非固定核。
- 实验验证对中等规模网络的有限宽度行为有准确描述,捕捉非高斯尾部等现象。
我们研究了宽贝叶斯神经网络,关注支配后验集中性的罕见但统计主导的波动,超越高斯过程极限。大偏差理论为预测器提供了显式的变分目标——率函数,从而在函数层面直接引入复杂度和特征学习的新概念。我们证明,后验输出率函数需通过预测器与内部核函数的联合优化得到,不同于固定核(NNGP)理论。数值实验表明,该方法能准确描述中等规模网络的有限宽度行为,捕捉非高斯尾部、后验形变及数据依赖的核选择效应。
原文摘要 · Abstract (English)
We study wide Bayesian neural networks focusing on the rare but statistically dominant fluctuations that govern posterior concentration, beyond Gaussian-process limits. Large-deviation theory provides explicit variational objectives-rate functions-on predictors, providing an emerging notion of complexity and feature learning directly at the functional level. We show that the posterior output rate function is obtained by a joint optimization over predictors and internal kernels, in contrast with fixed-kernel (NNGP) theory. Numerical experiments demonstrate that the resulting predictions accurately describe finite-width behavior for moderately sized networks, capturing non-Gaussian tails, posterior deformation, and data-dependent kernel selection effects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。