arXiv:2505.20172cs.LGmath.OC2025-05NeurIPS被引 11

小权重衰减让模型先快速拟合,再缓慢降低参数范数,解释了训练中突然泛化提升的现象。

A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation

  • 分两阶段:先沿无正则梯度流快速收敛到临界点流形
  • 后在1/λ时间尺度上缓慢减小参数的ℓ₂范数
  • 解释了深度学习中'突现泛化'现象,适合研究优化机制者阅读

我们研究了在一般损失函数 $F: \mathbb{R}^d \to \mathbb{R}$ 上,带小权重衰减 $λ$ 的梯度流动态。在温和正则性假设下,若无正则梯度流收敛,则当 $λ\to 0$ 时,轨迹呈现两阶段行为:初始快速阶段沿无正则梯度流移动,收敛至 $F$ 的临界点流形;随后在时间尺度 $1/λ$ 进入缓慢漂移阶段,遵循黎曼梯度流最小化参数的 $\ell_2$-范数。这一纯优化现象自然解释了深度学习中的'突现'(grokking)效应——训练损失迅速降至零,测试损失长期平稳后突然显著改善。我们认为该泛化跃迁源于权重衰减诱导的慢范数下降。我们在多个合成回归任务上验证了该机制。

原文摘要 · Abstract (English)

We study the dynamics of gradient flow with small weight decay on general training losses $F: \mathbb{R}^d \to \mathbb{R}$. Under mild regularity assumptions and assuming convergence of the unregularised gradient flow, we show that the trajectory with weight decay $λ$ exhibits a two-phase behaviour as $λ\to 0$. During the initial fast phase, the trajectory follows the unregularised gradient flow and converges to a manifold of critical points of $F$. Then, at time of order $1/λ$, the trajectory enters a slow drift phase and follows a Riemannian gradient flow minimising the $\ell_2$-norm of the parameters. This purely optimisation-based phenomenon offers a natural explanation for the \textit{grokking} effect observed in deep learning, where the training loss rapidly reaches zero while the test loss plateaus for an extended period before suddenly improving. We argue that this generalisation jump can be attributed to the slow norm reduction induced by weight decay, as explained by our analysis. We validate this mechanism empirically on several synthetic regression tasks.

优化理论泛化提升权重衰减

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。