无限宽贝叶斯神经网络在重尾先验下收敛到具有随机核的稳定过程。
Deep Kernel Posterior Learning under Infinite Variance Prior Weights
- 用重尾分布权重构建无限宽深层网络,使各层输出呈α-稳定分布
- 通过条件高斯表示实现递归连接各层协方差核,即使边际分布无定义
- 无需人工加噪,自然保留表征学习能力,适合理论与模型设计者
Neal(1996)证明,当权重先验方差有界时,无限宽浅层贝叶斯神经网络(BNN)收敛于高斯过程(GP)。Cho & Saul(2009)给出了深度核过程的递归公式,明确推导出多种常见激活函数下的逐层协方差核形式。然而,近期研究如Aitchison等(2021)指出,此类核为确定性,无法实现表征学习,即无法从数据中学习非退化的随机核后验。为此,他们引入人工噪声构造深度核逆威沙特过程,但该方法在经典BNN无限宽极限下缺乏自然基础。本文提出:若每层宽度趋于无穷,且所有权重服从椭球分布且方差无穷,则深层网络收敛至具有α-稳定边缘分布的过程,该过程具备条件高斯表示。尽管边际行为呈现稳定特性、协方差未必存在,仍可通过类似Cho & Saul的方式递归链接各层条件随机核。同时,我们推广了Loría & Bhadra(2024)关于浅层网络的最新结果至多层结构,并缓解其计算负担。模拟与基准数据集实验表明,该方法在计算与统计性能上显著优于现有方法。
原文摘要 · Abstract (English)
Neal (1996) proved that infinitely wide shallow Bayesian neural networks (BNN) converge to Gaussian processes (GP), when the network weights have bounded prior variance. Cho & Saul (2009) provided a useful recursive formula for deep kernel processes for relating the covariance kernel of each layer to the layer immediately below. Moreover, they worked out the form of the layer-wise covariance kernel in an explicit manner for several common activation functions. Recent works, including Aitchison et al. (2021), have highlighted that the covariance kernels obtained in this manner are deterministic and hence, precludes any possibility of representation learning, which amounts to learning a non-degenerate posterior of a random kernel given the data. To address this, they propose adding artificial noise to the kernel to retain stochasticity, and develop deep kernel inverse Wishart processes. Nonetheless, this artificial noise injection could be critiqued in that it would not naturally emerge in a classic BNN architecture under an infinite-width limit. To address this, we show that a Bayesian deep neural network, where each layer width approaches infinity, and all network weights are elliptically distributed with infinite variance, converges to a process with $α$-stable marginals in each layer that has a conditionally Gaussian representation. These conditional random covariance kernels could be recursively linked in the manner of Cho & Saul (2009), even though marginally the process exhibits stable behavior, and hence covariances are not even necessarily defined. We also provide useful generalizations of the recent results of Loría & Bhadra (2024) on shallow networks to multi-layer networks, and remedy the computational burden of their approach. The computational and statistical benefits over competing approaches stand out in simulations and in demonstrations on benchmark data sets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。