arXiv:2607.05735stat.MLcs.LG2026-07

证明了均值场贝叶斯神经网络在无限宽时仍保持有限宽的泛化能力。

Width-Robust Learnability in Mean-Field Bayesian Neural Networks

  • 通过缩减熵定义学习复杂度,建立宽度无关的可学习性判据。
  • 无限宽与多项式宽下,学习能力等价且由缩减熵是否多项式有界决定。
  • 适用于研究贝叶斯神经网络泛化机制的理论工作者。

无限宽极限是分析神经网络的常用方式,但其极限学习器未必具备大有限网络的复杂性归纳偏置。本文研究贝叶斯神经网络在均值场(临界特征学习)尺度下的情况。核心量为缩减熵: s_∞(y,ε) = ℎlman_N -\frac{1}{N}\log π_N^0(L ≤ ε), 即以群体均方误差ε表示目标函数y的先验代价强度。主要结果为宽度鲁棒可学习性定理:在固定深度下,布尔立方体类目标在无限宽时可由多项式样本学习,当且仅当在多项式宽时可学习,当且仅当其缩减熵多项式有界。换言之,忽略多项式精度损耗,贝叶斯均值场学习器仅在能被多项式大小网络表示的目标上实现精确泛化。正向证明采用子采样方法:从均值场解中无限多隐层神经元里,可选出多项式数量的代表性节点,仍能在所有输入上保持学习函数不变。在临界尺度下,该子采样包含‘主动’部分(保留数据依赖的低维统计),以及‘懒惰’部分(从先验重采样熵主导方向)。因此,无限宽均值场极限提供了无伪宽度依赖泛化能力的清晰解析描述。

原文摘要 · Abstract (English)

Infinite-width limits are a standard way to reason about neural networks, but it is not automatic that the limiting learner has the same complexity-theoretic inductive bias as large finite networks. We study this question for Bayesian neural networks at the mean-field, or critical feature-learning, scaling. The central quantity is the \emph{reduced entropy} \[ s_\infty(y,\varepsilon)=\limsup_N -\frac{1}{N}\log π_N^0(L\le \varepsilon), \] the intensive prior cost of representing a target function $y$ to population mean-squared error $\varepsilon$. Our main result is a width-robust learnability theorem. At fixed depth, a family of Boolean-cube targets is learnable from polynomially many samples at infinite width if and only if it is learnable at polynomial width, if and only if its reduced entropy is polynomially bounded. Equivalently, up to polynomial slack in accuracy, the Bayesian mean-field learner generalizes exactly on the targets that can be represented by polynomial-size networks. The forward direction is proved by a form of subsampling: from the infinitely many hidden neurons in the mean-field solution, one can select polynomially many representatives and still preserve the learned function on every input simultaneously. At the critical scaling this subsampling has both an ``active'' component, which keeps the data-dependent low-dimensional statistics, and a ``lazy'' component, which resamples the entropy-dominated directions from the prior. Thus the infinite-width mean-field limit gives a clean analytic description of learning without introducing spurious width-dependent generalization power.

贝叶斯神经网络泛化理论均值场可学习性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。