arXiv:2512.13227math.OCcs.LG2025-12被引 3

改进基于LMO的动量法,用二阶信息实现更快收敛

Better LMO-based Momentum Methods with Second-Order Information

  • 将海森修正动量融入LMO框架,适应任意范数
  • 理论证明收敛率提升至1/K^(1/3),优于传统方法的1/K^(1/4)
  • 适用于深层网络训练,适合关注优化效率的研究者

动量在随机优化算法中已被广泛证实有效。近期基于线性最小化算子(LMO)框架涌现出一批新方法,如Muon、Scion和Gluon,能高效解决深度神经网络训练问题。然而,传统随机动量方法的收敛速度理论上限仅为O(1/K^{1/4})。尽管已有方法如海森修正动量(HCM)试图提升该速率,其理论结果通常仅限于欧几里得范数设定,限制了在需任意范数问题中的应用。本文通过将HCM集成到LMO框架,实现了在放松光滑性假设与任意范数条件下的收敛性保证。我们建立了HCM的改进收敛率:O(1/K^{1/3}),可自适应问题几何结构,比传统动量更快。在多层感知机(MLP)和长短期记忆(LSTM)网络上的实验验证了理论发现。

原文摘要 · Abstract (English)

The use of momentum in stochastic optimization algorithms has shown empirical success across a range of machine learning tasks. Recently, a new class of stochastic momentum algorithms has emerged within the Linear Minimization Oracle (LMO) framework--leading to state-of-the-art methods, such as Muon, Scion, and Gluon, that effectively solve deep neural network training problems. However, traditional stochastic momentum methods offer convergence guarantees no better than the ${O}(1/K^{1/4})$ rate. While several approaches--such as Hessian-Corrected Momentum (HCM)--have aimed to improve this rate, their theoretical results are generally restricted to the Euclidean norm setting. This limitation hinders their applicability in problems, where arbitrary norms are often required. In this paper, we extend the LMO-based framework by integrating HCM, and provide convergence guarantees under relaxed smoothness and arbitrary norm settings. We establish improved convergence rates of ${O}(1/K^{1/3})$ for HCM, which can adapt to the geometry of the problem and achieve a faster rate than traditional momentum. Experimental results on training Multi-Layer Perceptrons (MLPs) and Long Short-Term Memory (LSTM) networks verify our theoretical observations.

优化算法动量方法二阶信息LMO框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。