提出新方法预测贝叶斯深度网络泛化性能,关键突破在有限宽度下捕捉核的层级重标度。
Kernel Renormalization in Bayesian Deep Neural Networks: the Equivalent Wishart Ansatz in the Proportional Regime

- 用等效威沙特假设建模深层网络核的随机波动,简化复杂性。
- 在比例极限下,仅需最多L个标量参数即可描述强表征学习效果。
- 适用于深度10、数据量千级的贝叶斯网络,可推广至卷积架构。
当训练样本数P与深度神经网络宽度N以相同比例增长(即比例宽度极限)时,浅层单隐层网络已有深入研究。但将此类非微扰结果扩展至深层非线性网络极为困难。本文提出一种有效近似方法,用于预测固定深度L的贝叶斯多层感知机(MLP)在任意高维数据上的泛化性能。我们引入等效威沙特假设,以捕捉MLP层级经验核的主导随机波动。该方法使我们在比例极限下对MLP的配分函数进行大偏差分析,其表达形式基于重标度的NNGP核。在此框架中,即使在比例极限下出现强烈表征学习,也仅通过最多L个标量序参数自洽确定。进一步将该方法扩展至卷积架构(CNNs),识别出一种层级局部核重标度机制,能量化因有限宽度效应导致的更复杂数据依赖型核变换。我们通过采样实验验证了该有效理论在深度约10、样本量约10³的有限深度神经网络贝叶斯后验上表现良好,整体吻合度高,并发现两类系统性偏差。
原文摘要 · Abstract (English)
The scaling limit where both the size of the training set $P$ and the width $N$ of a deep neural network grow at the same rate, the so-called proportional-width regime, has been intensely studied for shallow, single-hidden-layer networks. However, extending these non-perturbative results from shallow architectures to deep non-linear networks has proven very challenging. Here we present an effective approximate approach to predict the generalization performance of Bayesian multi-layer perceptrons (MLPs) of fixed depth $L$ on arbitrary high-dimensional data. We propose an equivalent Wishart Ansatz to capture the dominant stochastic fluctuations of the hierarchical empirical kernels of MLPs. This allows us to perform a large deviation analysis for the partition function of MLPs in the proportional limit, expressed in terms of a renormalized NNGP kernel. In this description, even strong representation learning in the proportional limit is encoded in at most $L$ scalar order parameters, determined self-consistently. Extending the approach to convolutional architectures (CNNs), we identify a hierarchical local kernel renormalization mechanism, which allows to quantify more complex data-dependent transformations of the large-width kernel in CNNs due to finite-width effects. We test our effective theory against sampling experiments from the Bayesian posterior of finite deep neural networks with depths $L \sim O(10)$ and $P\sim O(10^3)$ on classic benchmark datasets, finding overall very good agreement together with two distinct types of systematic deviations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。