arXiv:2501.18530stat.MLcond-mat.dis-nn2025-01被引 18

揭示浅层网络在插值附近的学习相变与最优泛化规律

Optimal generalisation and learning transition in extensive-width shallow neural networks near interpolation

  • 构建大宽度两层网络的贝叶斯最优泛化理论,支持任意激活函数
  • 发现泛化误差随采样率 $n/d^2$ 出现突变相变,分出普适与特化两阶段
  • 提出可预测但难优化的最优解,适合研究泛化极限的理论学者

我们研究了一个教师-学生监督学习模型,其中全训练的两层神经网络宽度 $k$ 和输入维度 $d$ 均较大且成比例。在样本量 $n$ 与输入维度平方成正比(即 $n \sim d^2$)的范围内,我们提供了适用于任意激活函数的贝叶斯最优泛化误差的有效理论,该范围接近插值阈值,此时可训练参数 $kd+k$ 与数据量 $n$ 相当。我们的分析处理了通用权重分布,揭示了一个从“普适”相到“特化”相的不连续相变。在普适相中,泛化误差独立于权重分布,随采样率 $n/d^2$ 缓慢下降,学生仅学习教师权重的非线性组合;在特化相中,误差依赖权重分布,且因学生与教师网络对齐而下降更快。因此,我们在插值附近发现了高度可预测的解,但可能难以被实际算法找到。

原文摘要 · Abstract (English)

We consider a teacher-student model of supervised learning with a fully-trained two-layer neural network whose width $k$ and input dimension $d$ are large and proportional. We provide an effective theory for approximating the Bayes-optimal generalisation error of the network for any activation function in the regime of sample size $n$ scaling quadratically with the input dimension, i.e., around the interpolation threshold where the number of trainable parameters $kd+k$ and of data $n$ are comparable. Our analysis tackles generic weight distributions. We uncover a discontinuous phase transition separating a "universal" phase from a "specialisation" phase. In the first, the generalisation error is independent of the weight distribution and decays slowly with the sampling rate $n/d^2$, with the student learning only some non-linear combinations of the teacher weights. In the latter, the error is weight distribution-dependent and decays faster due to the alignment of the student towards the teacher network. We thus unveil the existence of a highly predictive solution near interpolation, which is however potentially hard to find by practical algorithms.

泛化理论神经网络相变深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。