arXiv:2505.11748cs.LG2025-05

用三阶梯度构建动量,加速模型收敛并跳出局部最优

HOME-3: High-Order Momentum Estimator with Third-Power Gradient for Convex and Smooth Nonconvex Optimization

  • 用一阶梯度的立方构建高阶动量,突破传统平方动量限制
  • 理论证明三阶动量可改善凸与光滑非凸问题的收敛界
  • 在各类优化任务中均优于传统方法,适合追求高效优化的科研者

基于动量的梯度在优化先进机器学习模型中至关重要,不仅能加速收敛,还能帮助优化器逃离驻点。尽管当前主流动量方法多采用低阶梯度(如一阶梯度的平方),对更高阶梯度(幂次大于二)的研究仍较有限。本文提出高阶动量概念,聚焦于一阶梯度的立方作为代表性案例。理论上,我们证明引入三阶梯度可改善梯度优化器在凸与光滑非凸问题中的收敛界;实证上,通过在凸、光滑非凸及非光滑非凸优化任务上的广泛实验验证,高阶动量始终优于传统低阶动量方法,在多种优化场景中表现更优。

原文摘要 · Abstract (English)

Momentum-based gradients are essential for optimizing advanced machine learning models, as they not only accelerate convergence but also advance optimizers to escape stationary points. While most state-of-the-art momentum techniques utilize lower-order gradients, such as the squared first-order gradient, there has been limited exploration of higher-order gradients, particularly those raised to powers greater than two. In this work, we introduce the concept of high-order momentum, where momentum is constructed using higher-power gradients, with a focus on the third-power of the first-order gradient as a representative case. Our research offers both theoretical and empirical support for this approach. Theoretically, we demonstrate that incorporating third-power gradients can improve the convergence bounds of gradient-based optimizers for both convex and smooth nonconvex problems. Empirically, we validate these findings through extensive experiments across convex, smooth nonconvex, and nonsmooth nonconvex optimization tasks. Across all cases, high-order momentum consistently outperforms conventional low-order momentum methods, showcasing superior performance in various optimization problems.

优化算法动量方法高阶梯度收敛性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。