arXiv:2503.08489cs.LGcs.AI2025-03被引 1

提出三重惯性加速法,提升深度学习训练收敛速度与泛化能力

A Triple-Inertial Accelerated Alternating Optimization Method for Deep Learning Training

  • 采用三重惯性策略与专用近似方法,分项加速子问题优化
  • 在ReLU及其变体上显著减少迭代次数,提升训练效率
  • 理论证明全局收敛性与收敛速率,适合高维模型训练场景

随机梯度下降(SGD)在深度学习训练中表现优异,但存在梯度消失敏感、对输入数据敏感及缺乏可靠理论保证等问题。近年来,无梯度的交替最小化(AM)方法成为有潜力的替代方案,但收敛速度慢。为此,本文提出一种新型三重惯性加速交替最小化(TIAM)框架。TIAM通过三重惯性加速策略与专门设计的近似方法,实现对各子问题中不同项的针对性加速,显著提升收敛效率。此外,本文给出了TIAM算法的收敛性分析,包括全局收敛性与收敛速率。大量实验验证了TIAM的有效性,在ReLU及其变体上相比现有方法展现出更优的泛化能力与计算效率。

原文摘要 · Abstract (English)

The stochastic gradient descent (SGD) algorithm has achieved remarkable success in training deep learning models. However, it has several limitations, including susceptibility to vanishing gradients, sensitivity to input data, and a lack of robust theoretical guarantees. In recent years, alternating minimization (AM) methods have emerged as a promising alternative for model training by employing gradient-free approaches to iteratively update model parameters. Despite their potential, these methods often exhibit slow convergence rates. To address this challenge, we propose a novel Triple-Inertial Accelerated Alternating Minimization (TIAM) framework for neural network training. The TIAM approach incorporates a triple-inertial acceleration strategy with a specialized approximation method, facilitating targeted acceleration of different terms in each sub-problem optimization. This integration improves the efficiency of convergence, achieving superior performance with fewer iterations. Additionally, we provide a convergence analysis of the TIAM algorithm, including its global convergence properties and convergence rate. Extensive experiments validate the effectiveness of the TIAM method, showing significant improvements in generalization capability and computational efficiency compared to existing approaches, particularly when applied to the rectified linear unit (ReLU) and its variants.

深度学习优化算法收敛性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。