解析无调度优化方法的收敛性,揭示其高效原因
Understanding Schedule-Free Methods in Nonconvex Optimization: Rate Guarantees and Escaping Saddles
- 基于李雅普诺夫分析,建立连续时间模型研究收敛机制
- 证明其在非凸优化中达到一阶方法最优最坏情况收敛率
- 证明只需一次小扰动即可避开鞍点,适合复杂训练场景
Schedule-Free 方法因其无需设计和调参学习率调度器而受到广泛关注,同时在性能上可媲美甚至超越经过调优的优化器。尽管其在实践中表现优异,但在现代机器学习目标通常为非凸的场景下,其收敛理论仍基本未被探索。本文对标准形式下的 Schedule-Free 梯度下降与随机梯度下降(无辅助修改或限制条件)进行了最坏情况分析,针对光滑但可能非凸的目标函数。通过分析其对应的连续时间极限常微分方程所导出的李雅普诺夫函数,我们证明 Schedule-Free 梯度下降与随机梯度下降在所有一阶方法中达到了最优最坏情况收敛率。进一步地,我们将 Schedule-Free 梯度下降建模为非自治动力系统,并证明在任意小的一次性扰动下可实现严格鞍点规避。这些理论结果为 Schedule-Free 方法的强大性能提供了更深入的理解。
原文摘要 · Abstract (English)
Schedule-Free methods have attracted growing interest for alleviating the burden of designing and tuning a learning rate scheduler, while matching and sometimes even outperforming optimizers with tuned schedulers. Despite their strong empirical results, their convergence theory in nonconvex optimization, where modern machine learning objectives typically arise, has remained largely unexplored. In this paper, we provide worst-case analyses of Schedule-Free gradient descent and Schedule-Free stochastic gradient descent, in their standard form and without auxiliary modifications or restrictive conditions, for smooth but possibly nonconvex objectives. Based on a Lyapunov analysis derived from the continuous-time limiting ordinary differential equation associated with these methods, we show that Schedule-Free gradient descent and Schedule-Free stochastic gradient descent achieve the optimal worst-case convergence rates attainable among first-order methods. We further formulate Schedule-Free gradient descent as a nonautonomous dynamical system and prove strict-saddle avoidance under an arbitrarily small one-time perturbation. These theoretical results provide a better understanding of the strong performance that Schedule-Free methods demonstrate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。