用简单模型揭示大学习率下的稳定边界与渐进尖锐化机制
A Minimalist Example of Edge-of-Stability and Progressive Sharpening
- 构建二维输入的两层网络,解析大学习率下训练动态
- 证明训练全程存在渐进尖锐化与自稳现象,给出非渐近分析
- 连接极简模型与通用理论,为优化策略提供新视角
深度学习优化近年揭示了大学习率下的两个有趣现象:稳定边界(Edge of Stability, EoS)和渐进尖锐化(Progressive Sharpening, PS),挑战了经典梯度下降(GD)分析。现有研究或采用泛化框架,或依赖极简示例,均面临解释力不足的问题。本文通过引入一个两层网络,其输入为二维:一维相关、一维无关,严格证明了在大学习率下存在渐进尖锐化与自稳现象,并对整个梯度下降轨迹上的训练动态与尖锐度进行了非渐近分析。此外,我们通过重建极简模型与泛化分析间的‘稳定集’,将梯度流解的尖锐度分析扩展至二维输入场景。这些发现从参数与数据分布双重视角深化了对EoS的理解,可能为实际深度学习优化提供更有效的策略。
原文摘要 · Abstract (English)
Recent advances in deep learning optimization have unveiled two intriguing phenomena under large learning rates: Edge of Stability (EoS) and Progressive Sharpening (PS), challenging classical Gradient Descent (GD) analyses. Current research approaches, using either generalist frameworks or minimalist examples, face significant limitations in explaining these phenomena. This paper advances the minimalist approach by introducing a two-layer network with a two-dimensional input, where one dimension is relevant to the response and the other is irrelevant. Through this model, we rigorously prove the existence of progressive sharpening and self-stabilization under large learning rates, and establish non-asymptotic analysis of the training dynamics and sharpness along the entire GD trajectory. Besides, we connect our minimalist example to existing works by reconciling the existence of a well-behaved ``stable set" between minimalist and generalist analyses, and extending the analysis of Gradient Flow Solution sharpness to our two-dimensional input scenario. These findings provide new insights into the EoS phenomenon from both parameter and input data distribution perspectives, potentially informing more effective optimization strategies in deep learning practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。