arXiv:2410.24206cs.LGcs.AI2024-10被引 57

用平均轨迹建模深度学习优化,揭示震荡中的进步机制

Understanding Optimization in Deep Learning with Central Flows

  • 提出中心流理论,通过时间平均轨迹分析震荡优化过程
  • 实验证明中心流能高精度预测任意神经网络的长期优化路径
  • 解释自适应优化器如何隐式放大步长,突破传统理解

传统优化理论无法描述深度学习中优化过程的动力学,即使在确定性训练的简单情形下亦然。挑战在于优化器通常处于一种复杂的振荡状态,称为“稳定边缘”。本文提出可描述该状态下优化动力学的新理论。核心洞察是:尽管振荡优化器的精确轨迹难以分析,但其时间平均(即平滑)轨迹往往更易处理。为此,我们推导出一种称为“中心流”的微分方程,刻画这一时间平均轨迹。实验表明,这些中心流可对任意神经网络的长期优化轨迹实现高度数值精度的预测。通过解析中心流,我们揭示了梯度下降为何能在损失偶尔上升时仍持续进展;解释了自适应优化器如何“适应”局部损失曲面;并阐明了它们如何隐式导航至可采取更大步长的区域。结果表明,中心流可成为理解深度学习优化的重要理论工具。

原文摘要 · Abstract (English)

Traditional theories of optimization cannot describe the dynamics of optimization in deep learning, even in the simple setting of deterministic training. The challenge is that optimizers typically operate in a complex, oscillatory regime called the "edge of stability." In this paper, we develop theory that can describe the dynamics of optimization in this regime. Our key insight is that while the *exact* trajectory of an oscillatory optimizer may be challenging to analyze, the *time-averaged* (i.e. smoothed) trajectory is often much more tractable. To analyze an optimizer, we derive a differential equation called a "central flow" that characterizes this time-averaged trajectory. We empirically show that these central flows can predict long-term optimization trajectories for generic neural networks with a high degree of numerical accuracy. By interpreting these central flows, we are able to understand how gradient descent makes progress even as the loss sometimes goes up; how adaptive optimizers "adapt" to the local loss landscape; and how adaptive optimizers implicitly navigate towards regions where they can take larger steps. Our results suggest that central flows can be a valuable theoretical tool for reasoning about optimization in deep learning.

优化理论深度学习中心流

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。