提出新稳定边界理论,揭示优化器如何动态调节训练稳定性。
The Road Taken: The Role of Optimizers at the Edge of Stability
- 基于方向性海森值与梯度对齐度重构稳定阈值
- 发现传统理论低估实际稳定性达21.1倍且依赖优化器类型
- 提供诊断工具,揭示优化器在时空预算间的主动平衡作用
边缘稳定性指深度学习中基于梯度的优化器使损失函数的海森值特征值维持在经典下降引理预测的不稳定性阈值之上。以往研究以最大海森特征值和学习率定义该现象,但本文发现包括梯度下降在内的多种一阶优化方法实际违反此稳定性边界达21.1倍之多。该偏差具有系统性且高度依赖具体优化器,现有理论无法捕捉。为此,本文从优化器实际更新方向的定向海森值与梯度对齐度出发,推导出新的‘实现边缘稳定性’定义。新框架不仅消除了优化器相关的偏移,提升了稳定性阈值预测一致性,还引入新诊断工具,揭示优化器在第一阶优化中主动平衡时间与空间资源的机制。
原文摘要 · Abstract (English)
The edge of stability refers to a phenomenon in deep learning with gradient-based optimizers where the Hessian eigenvalues of the loss remain stable above a threshold that the classical descent lemma predicts to be unstable. Previous works formulate the edge of stability with respect to the maximum Hessian eigenvalue and the learning rate. However, we observe that many first-order methods, including gradient descent, significantly violate the stability bound predicted by these theories by a factor as large as $\times 21.1$. Moreover, this deviation turns out to be systematic and highly dependent on the underlying optimizer, which is not captured by previous formulations. This calls for a new formulation of the stability threshold, which we derive from the directional Hessian and the gradient-alignment score with respect to the actual update taken by the optimizer, rather than the maximum curvature mode. Our new formulation of the realized edge of stability not only removes optimizer-dependent offsets and provides more consistent predictions of the stability threshold, but also introduces new diagnostic tools that reveal the unique role of the optimizer in actively balancing between the temporal and spatial budgets in first-order optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。