arXiv:2603.28964cs.LGcs.AI2026-03被引 5

神经网络训练中的突变现象由参数更新的谱间隙决定,可预测学习行为转折点。

Spectral Edge Dynamics: An Analytical-Empirical Study of Phase Transitions in Neural Network Training

  • 用滚动窗口的格拉姆矩阵谱隙分析训练过程,揭示相变机制。
  • 谱隙位置 $k^*$ 是学习动态的关键,其崩溃会中断训练,且与优化器有关。
  • 适用于理解模型学习、调参和解释复杂训练现象的研究者。

我们提出谱边缘分析:神经网络训练中的相变现象——如突现学习、能力提升、损失平台期——受参数更新滚动窗口格拉姆矩阵谱隙的控制。在极端参数比($P ilde{10}^8$,窗口 $W ilde{10}$)下,经典BBP检测阈值失效;起作用的是主导模式与次主导模式之间的内信号隙,位于 $k^* = \mathrm{argmax}\, σ_j/σ_{j+1}$。基于三项假设推导出:(i) 谱隙动态遵循带有曲率不对称性、阻尼和梯度驱动的狄松型常微分方程;(ii) 谱损失分解将各模式的学习贡献与Davis--Kahan稳定性系数关联;(iii) 谱隙最大性原理表明,$k^*$ 是唯一动态特权位置——其坍塌是唯一破坏学习的事件,且通过无需优化器假设的 $α$-反馈环维持自身。绝热参数 $\mathcal{A} = \|ΔG\|_F / (η\, g^2)$ 控制电路稳定性:$\mathcal{A} \ll 1$(平台期),$\mathcal{A} \sim 1$(相变),$\mathcal{A} \gg 1$(遗忘)。在六类模型(15万至1.24亿参数)上验证:谱隙动态先于所有突现事件(24/24有权重衰减,1/24无),谱隙位置随优化器变化(同一模型下,Muon: $k^*=1$,AdamW: $k^*=2$),19/20定量预测被证实。该框架与边缘稳定、张量程序、狄松布朗运动、彩票假说及神经网络缩放定律一致。

原文摘要 · Abstract (English)

We develop the spectral edge analysis: phase transitions in neural network training -- grokking, capability gains, loss plateaus -- are controlled by the spectral gap of the rolling-window Gram matrix of parameter updates. In the extreme aspect ratio regime (parameters $P \sim 10^8$, window $W \sim 10$), the classical BBP detection threshold is vacuous; the operative structure is the intra-signal gap separating dominant from subdominant modes at position $k^* = \mathrm{argmax}\, σ_j/σ_{j+1}$. From three assumptions we derive: (i) gap dynamics governed by a Dyson-type ODE with curvature asymmetry, damping, and gradient driving; (ii) a spectral loss decomposition linking each mode's learning contribution to its Davis--Kahan stability coefficient; (iii) the Gap Maximality Principle, showing that $k^*$ is the unique dynamically privileged position -- its collapse is the only one that disrupts learning, and it sustains itself through an $α$-feedback loop requiring no assumption on the optimizer. The adiabatic parameter $\mathcal{A} = \|ΔG\|_F / (η\, g^2)$ controls circuit stability: $\mathcal{A} \ll 1$ (plateau), $\mathcal{A} \sim 1$ (phase transition), $\mathcal{A} \gg 1$ (forgetting). Tested across six model families (150K--124M parameters): gap dynamics precede every grokking event (24/24 with weight decay, 1/24 without), the gap position is optimizer-dependent (Muon: $k^*=1$, AdamW: $k^*=2$ on the same model), and 19/20 quantitative predictions are confirmed. The framework is consistent with the edge of stability, Tensor Programs, Dyson Brownian motion, the Lottery Ticket Hypothesis, and neural scaling laws.

神经网络训练相变谱分析学习动力学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。