arXiv:2607.04993cs.LGcs.AI2026-07被引 1

用动力系统视角揭示梯度下降的有限步行为如何决定模型学习路径。

The Map Behind the Flow: Finite-Step Gradient Descent as a Dynamical System

论文配图:The Map Behind the Flow: Finite-Step Gradient Descent as a Dynamical System
图 1 · 摘自论文原文
  • 将固定步长梯度下降视为离散动力系统,构建可解的深层线性模型。
  • 发现边缘稳定性实为训练映射的首次分岔,非优化失效。
  • 学习率是决定模型表示选择的结构性参数,影响平衡与平坦化。

深度学习中的许多现象具有动态特性:不仅关注极小值的存在,更关注梯度下降如何到达、避开或选择它们。边缘稳定性行为、尖锐度振荡、弹射阶段、平衡机制及向更平坦表示的演化,均源于训练映射本身的特性,无法用小步长梯度流近似捕捉。本文将固定步长梯度下降建模为离散动力系统,研究一系列精确可解的模型,保留深度、因子分解、宽度、数据耦合、激活函数和随机性等深度学习基本结构。起点为深层线性链的平衡标量简化,产生四次损失和三次梯度映射,其边缘后行为可显式表达。在自然的大深度缩放下,该动力学收敛至普适的Ricker型映射。因此,边缘稳定性并非优化崩溃,而是训练映射的首个分岔。将标量动力学嵌回因子模型,使这些区间转化为学习现象。有限步长破坏梯度流的守恒律,收缩因子不平衡;残余振荡推动参数向更平坦、更平衡的表示移动。更宽的线性网络产生一系列谱边,最优学习率可能位于首条边之外。数据耦合、非线性激活和随机目标保持相同组织原则:有限步长振荡驱动对齐、平衡与表征选择。因此,学习率不仅是数值稳定性参数,更是训练动力学的结构性参数,决定其吸引子并塑造梯度下降所选的表征。

原文摘要 · Abstract (English)

Many phenomena of deep learning are dynamical: they concern not only which minima exist, but how gradient descent reaches, avoids, or selects among them. Edge-of-stability behavior, sharpness oscillations, catapult phases, balancing, and movement toward flatter representations are effects of the training map itself, and are poorly captured by the small-step gradient-flow limit. This paper studies fixed-step gradient descent as a discrete dynamical system in a hierarchy of exactly solvable models retaining basic structures of deep learning: depth, factorization, width, data coupling, activation, and stochasticity. The starting point is the balanced scalar reduction of a deep linear chain, giving a quartic loss and a cubic gradient map whose post-edge behavior is explicit. Under the natural large-depth scaling, this dynamics converges to a universal Ricker-type map. The edge of stability is therefore not a breakdown of optimization, but the first bifurcation of the training map. Embedding the scalar dynamics back into factored models turns these regimes into learning phenomena. Finite steps break conservation laws of gradient flow and contract factorization imbalance; residual oscillations move parameters toward flatter, more balanced representations. Wider linear networks produce a ladder of spectral edges, so the optimal learning rate can lie beyond the first edge. Data coupling, nonlinear activations, and stochastic targets preserve the same organizing principle: finite-step oscillations drive alignment, balancing, and representation selection. Thus the learning rate is not merely a numerical stability parameter. It is a structural parameter of the training dynamics, determining its attractors and shaping the representations gradient descent selects.

动力系统梯度下降学习率表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。