揭示神经网络优化中边缘稳定性边界的非凸性突破机制
Conservation Law Breaking at the Edge of Stability: A Spectral Theory of Non-Convex Neural Network Optimization
- 发现梯度流下存在保结构的L-1守恒律,约束优化轨迹于低维流形
- 离散梯度下降导致守恒律破坏,漂移量与学习率幂次相关(约1.1-1.6)
- 提出可解释的谱交叉公式,适用于线性与ReLU网络,适合研究优化动力学者
为何梯度下降在非凸神经网络优化中仍能稳定找到优质解,尽管最坏情况下问题属于NP难?本文发现,无偏置的L层ReLU网络在梯度流下保持L-1守恒律:C_l = ||W_{l+1}||_F^2 - ||W_l||_F^2,使优化轨迹被限制在低维流形上。在离散梯度下降中,该守恒律被打破,总漂移量随学习率η的幂次α变化,α值约为1.1至1.6,具体取决于架构、损失函数和网络宽度。漂移可精确分解为η²·S(η),其中梯度失衡和S(η)具有闭式谱交叉公式,模式系数c_k正比于e_k(0)²·λ_{x,k}²,由第一性原理推导并验证于线性网络(R=0.85)与ReLU网络(R>0.80)。对于交叉熵损失,软最大概率集中驱动赫斯矩阵谱指数压缩,时间尺度τ = Θ(1/η),与训练集大小无关,解释了为何交叉熵使漂移指数自正则化至α≈1.0。识别出两种动力学相变:由宽度决定的过渡点,一侧为微扰子稳定区(谱公式适用),另一侧为强耦合非微扰区。所有预测在23个实验中得到验证。
原文摘要 · Abstract (English)
Why does gradient descent reliably find good solutions in non-convex neural network optimization, despite the landscape being NP-hard in the worst case? We show that gradient flow on L-layer ReLU networks without bias preserves L-1 conservation laws C_l = ||W_{l+1}||_F^2 - ||W_l||_F^2, confining trajectories to lower-dimensional manifolds. Under discrete gradient descent, these laws break with total drift scaling as eta^alpha where alpha is approximately 1.1-1.6 depending on architecture, loss function, and width. We decompose this drift exactly as eta^2 * S(eta), where the gradient imbalance sum S(eta) admits a closed-form spectral crossover formula with mode coefficients c_k proportional to e_k(0)^2 * lambda_{x,k}^2, derived from first principles and validated for both linear (R=0.85) and ReLU (R>0.80) networks. For cross-entropy loss, softmax probability concentration drives exponential Hessian spectral compression with timescale tau = Theta(1/eta) independent of training set size, explaining why cross-entropy self-regularizes the drift exponent near alpha=1.0. We identify two dynamical regimes separated by a width-dependent transition: a perturbative sub-Edge-of-Stability regime where the spectral formula applies, and a non-perturbative regime with extensive mode coupling. All predictions are validated across 23 experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。