arXiv:2504.19983cs.LGstat.ML2025-04NeurIPS被引 33

揭示浅层神经网络在SGD训练中的涌现规律与缩放法则。

Emergence and scaling laws in SGD learning of shallow neural networks

  • 通过精确分析揭示学习过程中信号方向的突变式恢复时间。
  • 发现损失函数随样本数、迭代步数和参数量的幂律缩放关系。
  • 适用于研究大规模神经网络训练动力学与泛化性能的理论分析。

我们研究了在各向同性高斯数据下,对含有P个神经元的两层神经网络进行在线随机梯度下降(SGD)学习的复杂性:目标函数为 $f_*(oldsymbol{x}) = \ sum_{p=1}^P a_p\cdot σ(\langle\boldsymbol{x},\boldsymbol{v}_p^*\rangle)$,其中 $\boldsymbol{x} \sim \mathcal{N}(0,\boldsymbol{I}_d)$,激活函数 $σ$ 为偶函数且信息指数 $k_*>2$(定义为赫尔米特展开中的最低阶次),$\\\{oldsymbol{v}^*_p\\}_{p\in[P]}\subset \mathbb{R}^d$ 为正交信号方向,第二层系数非负且满足 $\sum_{p} a_p^2=1$。研究聚焦于 $P\gg 1$ 的广义宽度情形,并允许第二层条件数发散,涵盖幂律情形 $a_p\asymp p^{-β}$($β\in\mathbb{R}_{\ge 0}$)。我们对学生网络最小化均方误差(MSE)目标下的训练动态进行了精确分析,明确识别出每个信号方向恢复的尖锐过渡时间。在幂律设定下,刻画了损失函数关于训练样本数、SGD步数及学生网络参数量的缩放指数。分析表明,尽管单个教师神经元的学习呈现突变特征,但 $P\gg 1$ 个神经元在不同时间尺度上的并行学习导致累积目标函数出现平滑的缩放规律。

原文摘要 · Abstract (English)

We study the complexity of online stochastic gradient descent (SGD) for learning a two-layer neural network with $P$ neurons on isotropic Gaussian data: $f_*(\boldsymbol{x}) = \sum_{p=1}^P a_p\cdot σ(\langle\boldsymbol{x},\boldsymbol{v}_p^*\rangle)$, $\boldsymbol{x} \sim \mathcal{N}(0,\boldsymbol{I}_d)$, where the activation $σ:\mathbb{R}\to\mathbb{R}$ is an even function with information exponent $k_*>2$ (defined as the lowest degree in the Hermite expansion), $\{\boldsymbol{v}^*_p\}_{p\in[P]}\subset \mathbb{R}^d$ are orthonormal signal directions, and the non-negative second-layer coefficients satisfy $\sum_{p} a_p^2=1$. We focus on the challenging ``extensive-width'' regime $P\gg 1$ and permit diverging condition number in the second-layer, covering as a special case the power-law scaling $a_p\asymp p^{-β}$ where $β\in\mathbb{R}_{\ge 0}$. We provide a precise analysis of SGD dynamics for the training of a student two-layer network to minimize the mean squared error (MSE) objective, and explicitly identify sharp transition times to recover each signal direction. In the power-law setting, we characterize scaling law exponents for the MSE loss with respect to the number of training samples and SGD steps, as well as the number of parameters in the student neural network. Our analysis entails that while the learning of individual teacher neurons exhibits abrupt transitions, the juxtaposition of $P\gg 1$ emergent learning curves at different timescales leads to a smooth scaling law in the cumulative objective.

神经网络SGD缩放律学习动力学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。