arXiv:2604.19740cs.LGcs.AI2026-04

揭示大学习率下模型泛化背后的几何机制

Generalization at the Edge of Stability

论文配图:Generalization at the Edge of Stability
图 1 · 摘自论文原文
  • 将优化器视为随机动力系统,发现其收敛到低维分形吸引子
  • 提出'尖锐维度'新度量,证明其与泛化误差直接相关
  • 解释了混沌训练中模型为何能更好泛化,适合研究优化理论者

现代神经网络训练常采用大学习率,处于不稳定的边缘,此时优化过程呈现振荡和混沌行为。实证表明该区域往往具有更好的泛化性能,但机理尚不清晰。本文将随机优化器建模为随机动力系统,发现其通常收敛至分形吸引子集,且内在维数较低。基于李雅普诺夫维数理论,我们引入新概念‘尖锐维度’,并建立基于此的泛化界。结果表明,混沌状态下的泛化能力依赖于完整海森矩阵谱及部分行列式的结构,这是以往仅用迹或谱范数无法捕捉的复杂性。在多种MLP和Transformer上的实验验证了理论,并为近期观测到的‘顿悟现象’提供了新解释。

原文摘要 · Abstract (English)

Training modern neural networks often relies on large learning rates, operating at the edge of stability, where the optimization dynamics exhibit oscillatory and chaotic behavior. Empirically, this regime often yields improved generalization performance, yet the underlying mechanism remains poorly understood. In this work, we represent stochastic optimizers as random dynamical systems, which often converge to a fractal attractor set (rather than a point) with a smaller intrinsic dimension. Building on this connection and inspired by Lyapunov dimension theory, we introduce a novel notion of dimension, coined the `sharpness dimension', and prove a generalization bound based on this dimension. Our results show that generalization in the chaotic regime depends on the complete Hessian spectrum and the structure of its partial determinants, highlighting a complexity that cannot be captured by the trace or spectral norm considered in prior work. Experiments across various MLPs and transformers validate our theory while also providing new insights into the recently observed phenomenon of grokking.

优化理论泛化分析混沌训练尖锐维度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。