arXiv:2512.15606stat.MLcs.LG2025-12

研究神经网络接近最优时的学习动态,揭示海森矩阵决定学习性能的关键作用。

A Teacher-Student Perspective on the Dynamics of Learning Near the Optimal Point

  • 通过师生模型分析海森矩阵谱分布,发现小特征值主导长期学习表现。
  • 线性网络的海森谱渐近服从缩放卡方与马尔琴科-帕斯图尔分布的卷积。
  • 多项式激活函数下海森秩反映有效参数量,非线性激活则保持满秩。

当神经网络接近最优学习点时,梯度下降的性能由损失函数关于网络参数的海森矩阵决定。本文针对教师-学生问题中权重匹配的情形,刻画了特定类别的海森矩阵特征谱,表明其较小特征值决定长期学习表现。对于线性网络,我们严格证明大网络下谱分布渐近为缩放后的卡方分布与缩放后的马尔琴科-帕斯图尔分布的卷积。数值上分析了多项式及其他非线性网络的海森谱。此外,我们发现多项式激活函数网络的海森矩阵秩可视为有效参数数量;而对于如误差函数等通用非线性激活函数,经验上始终为满秩。

原文摘要 · Abstract (English)

Near an optimal learning point of a neural network, the learning performance of gradient descent dynamics is dictated by the Hessian matrix of the loss function with respect to the network parameters. We characterize the Hessian eigenspectrum for some classes of teacher-student problems, when the teacher and student networks have matching weights, showing that the smaller eigenvalues of the Hessian determine long-time learning performance. For linear networks, we analytically establish that for large networks the spectrum asymptotically follows a convolution of a scaled chi-square distribution with a scaled Marchenko-Pastur distribution. We numerically analyse the Hessian spectrum for polynomial and other non-linear networks. Furthermore, we show that the rank of the Hessian matrix can be seen as an effective number of parameters for networks using polynomial activation functions. For a generic non-linear activation function, such as the error function, we empirically observe that the Hessian matrix is always full rank.

神经网络海森矩阵学习动态谱分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。