揭示浅层网络在插值附近的学习相变与最优泛化规律
Optimal generalisation and learning transition in extensive-width shallow neural networks near interpolation
- 构建大宽度两层网络的贝叶斯最优泛化理论,支持任意激活函数
- 发现泛化误差随采样率 $n/d^2$ 出现突变相变,分出普适与特化两阶段
- 提出可预测但难优化的最优解,适合研究泛化极限的理论学者
我们研究了一个教师-学生监督学习模型,其中全训练的两层神经网络宽度 $k$ 和输入维度 $d$ 均较大且成比例。在样本量 $n$ 与输入维度平方成正比(即 $n \sim d^2$)的范围内,我们提供了适用于任意激活函数的贝叶斯最优泛化误差的有效理论,该范围接近插值阈值,此时可训练参数 $kd+k$ 与数据量 $n$ 相当。我们的分析处理了通用权重分布,揭示了一个从“普适”相到“特化”相的不连续相变。在普适相中,泛化误差独立于权重分布,随采样率 $n/d^2$ 缓慢下降,学生仅学习教师权重的非线性组合;在特化相中,误差依赖权重分布,且因学生与教师网络对齐而下降更快。因此,我们在插值附近发现了高度可预测的解,但可能难以被实际算法找到。
原文摘要 · Abstract (English)
We consider a teacher-student model of supervised learning with a fully-trained two-layer neural network whose width $k$ and input dimension $d$ are large and proportional. We provide an effective theory for approximating the Bayes-optimal generalisation error of the network for any activation function in the regime of sample size $n$ scaling quadratically with the input dimension, i.e., around the interpolation threshold where the number of trainable parameters $kd+k$ and of data $n$ are comparable. Our analysis tackles generic weight distributions. We uncover a discontinuous phase transition separating a "universal" phase from a "specialisation" phase. In the first, the generalisation error is independent of the weight distribution and decays slowly with the sampling rate $n/d^2$, with the student learning only some non-linear combinations of the teacher weights. In the latter, the error is weight distribution-dependent and decays faster due to the alignment of the student towards the teacher network. We thus unveil the existence of a highly predictive solution near interpolation, which is however potentially hard to find by practical algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。