揭示神经网络训练中谱相变如何影响可训练性
Spectral phase transitions and trainability in neural network learning dynamics

- 将训练视为随机矩阵的演化,用梯度下降重塑谱结构
- 发现学习率与初始权重方差决定训练成败的相图
- 适用于理解模型泛化与表示学习的统一框架
神经网络权重矩阵谱中低维结构的出现是训练后模型的常见现象,但其动态起源仍不明确。我们将神经网络训练建模为由随机梯度下降驱动的初始随机矩阵集合的随机演化,该过程重塑谱体并增强信号强度,引发贝克-本·阿鲁斯-佩谢(BBP)相变:孤立特征值从随机谱体中分离,为高维学习动态中的表征形成提供了动力学框架。我们在一个可解析求解的线性教师-学生模型中验证了这一机制,得到了由步长(或学习率)和初始权重方差决定的可训练性相图,并将该形式拓展至非线性和随机设置。真实场景的数值模拟支持该观点,显示训练过程中谱对齐的稳健出现。结果表明,谱分析可能为随机学习动态提供统一视角,连接可训练性、优化超参数、谱相变与表示学习。
原文摘要 · Abstract (English)
The emergence of low-dimensional structures in the spectra of neural network weight matrices is a common empirical feature of trained models, but the dynamical origin of this phenomenon during learning remains an open problem. We formulate neural network training as the stochastic evolution of an initially random matrix ensemble, driven by stochastic gradient descent (SGD) updates that reshape the spectral bulk while amplifying signal strength. This induces a Baik-Ben Arous-Péché (BBP) transition during training, where isolated eigenvalues detach from the random bulk distribution, providing a dynamical framework for representation formation in high-dimensional learning dynamics. We demonstrate this in a solvable linear teacher-student model, where spectral evolution is analytically tractable and a phase diagram of trainability governed by the step size (or learning rate) and initial weight variance is obtained, and subsequently extend our formalism beyond the linear regime to nonlinear and stochastic settings. Numerical simulations in realistic settings support this picture, showing robust emergence of spectral alignment during training. Our results suggest that spectral analysis may provide a unified perspective of stochastic learning dynamics, linking trainability, optimisation hyperparameters, spectral phase transitions, and representation learning in neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。