arXiv:2501.09137cs.LGmath.OC2025-01被引 5

GD比梯度流收敛到更平坦的解,且步长越大越有隐式正则化。

Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks

  • GD在大步长下仍线性收敛,步长可达约2/尖锐度
  • 收敛解的范数和尖锐度低于梯度流解
  • 解释了'稳定边缘'训练的优势,适合研究优化机制

我们研究了深度为2、单输入单输出的线性神经网络中梯度下降(GD)的动力学。即使使用较大步长(约2/尖锐度),GD仍能以显式线性速率收敛至全局最小值;步长更大时虽仍收敛,但速度极慢。我们还刻画了GD最终收敛的解:其权重范数和尖锐度均低于梯度流解。分析揭示了收敛速度与隐式正则化强度之间的权衡,解释了在'稳定边缘'训练带来的额外正则化效果,对复杂模型训练具有启示意义。

原文摘要 · Abstract (English)

We study the gradient descent (GD) dynamics of a depth-2 linear neural network with a single input and output. We show that GD converges at an explicit linear rate to a global minimum of the training loss, even with a large stepsize -- about $2/\textrm{sharpness}$. It still converges for even larger stepsizes, but may do so very slowly. We also characterize the solution to which GD converges, which has lower norm and sharpness than the gradient flow solution. Our analysis reveals a trade off between the speed of convergence and the magnitude of implicit regularization. This sheds light on the benefits of training at the ``Edge of Stability'', which induces additional regularization by delaying convergence and may have implications for training more complex models.

优化理论梯度下降隐式正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。