arXiv:2411.07061cs.LGmath.OC2024-11被引 13

证明了无需学习率调度的SGD在非凸优化中同样高效

General framework for online-to-nonconvex conversion: Schedule-free SGD is also effective for nonconvex optimization

  • 提出在线学习转非凸优化的通用框架
  • 证明无调度SGD在非凸问题上达到最优迭代复杂度
  • 解释参数选择,填补理论空白,适合优化研究者

本文研究了由A. Defazio等人(NeurIPS 2024)提出的无调度方法在非凸优化中的有效性,其在训练神经网络中表现出色。我们证明,无调度SGD在非光滑、非凸优化问题上达到了最优迭代复杂度。我们的分析建立了一个通用的在线学习到非凸优化的转换框架,该框架不仅能复现已有转换方法,还引出了两种新转换方案。其中一种直接对应于无调度SGD,从而使其最优性得以确立。此外,我们的分析为无调度SGD的参数选择提供了理论洞见,弥补了凸优化理论无法解释的空白。

原文摘要 · Abstract (English)

This work investigates the effectiveness of schedule-free methods, developed by A. Defazio et al. (NeurIPS 2024), in nonconvex optimization settings, inspired by their remarkable empirical success in training neural networks. Specifically, we show that schedule-free SGD achieves optimal iteration complexity for nonsmooth, nonconvex optimization problems. Our proof begins with the development of a general framework for online-to-nonconvex conversion, which converts a given online learning algorithm into an optimization algorithm for nonconvex losses. Our general framework not only recovers existing conversions but also leads to two novel conversion schemes. Notably, one of these new conversions corresponds directly to schedule-free SGD, allowing us to establish its optimality. Additionally, our analysis provides valuable insights into the parameter choices for schedule-free SGD, addressing a theoretical gap that the convex theory cannot explain.

优化算法非凸优化无调度SGD理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。