arXiv:2505.21722cs.LGcs.AI2025-05中稿 · ICLR被引 4

揭示深度ReLU网络梯度下降的鞍点逃逸路径具有深层低秩特性。

Saddle-To-Saddle Dynamics in Deep ReLU Networks: Low-Rank Bias in the First Saddle Escape

  • 分析梯度下降在参数空间原点鞍点处的逃逸方向
  • 发现深层权重矩阵首奇异值比其他值大至少ℓ^(1/4)倍
  • 为理解深层网络优化动态提供新视角,适合研究优化机制者

当深度ReLU网络以小权重初始化时,梯度下降最初受参数空间原点鞍点主导。本文研究了梯度下降离开原点的逃逸方向,这些方向在作用上类似于严格鞍点的海森矩阵特征向量。我们证明最优逃逸方向在深层具有显著低秩偏置:第ℓ层权重矩阵的首奇异值至少比其他奇异值大ℓ^(1/4)倍。此外,我们还推导出一系列关于这些逃逸方向的相关结论。研究建议深度ReLU网络呈现鞍点到鞍点的动力学行为,梯度下降依次经过瓶颈秩递增的多个鞍点(Jacot, 2023)。

原文摘要 · Abstract (English)

When a deep ReLU network is initialized with small weights, gradient descent (GD) is at first dominated by the saddle at the origin in parameter space. We study the so-called escape directions along which GD leaves the origin, which play a similar role as the eigenvectors of the Hessian for strict saddles. We show that the optimal escape direction features a low-rank bias in its deeper layers: the first singular value of the $\ell$-th layer weight matrix is at least $\ell^{\frac{1}{4}}$ larger than any other singular value. We also prove a number of related results about these escape directions. We suggest that deep ReLU networks exhibit saddle-to-saddle dynamics, with GD visiting a sequence of saddles with increasing bottleneck rank (Jacot, 2023).

深度学习优化动力学低秩结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。