arXiv:2410.09943cs.LGmath.OC2024-10

用非线性自回归模型动态调节学习率和动量,提升训练稳定性与收敛速度。

Dynamic Estimation of Learning Rates Using a Non-Linear Autoregressive Model

  • 基于带动量的非线性自回归框架,动态估计学习率与动量参数。
  • 在多个数据集和强化学习环境上验证,对大初始学习率有强鲁棒性。
  • 适合追求快速收敛与高稳定性的深度学习优化场景。

我们提出一类新型自适应非线性自回归(Nlar)模型,融合动量思想,随迭代次数动态估计学习率与动量。方法通过缩放(截断)函数控制梯度增长,实现稳定收敛。在此框架下,我们设计了三种学习率估计算法,并提供收敛性理论证明。进一步展示了这些估计算法如何支撑有效Nlar优化器的构建。通过在多个数据集及强化学习环境中的广泛实验,验证了所提算法与优化器的性能。结果表明,Nlar优化器具有两大特点:在底层参数变化(包括大初始学习率)下仍能稳健收敛;在初始阶段表现出强适应性与快速收敛能力。

原文摘要 · Abstract (English)

We introduce a new class of adaptive non-linear autoregressive (Nlar) models incorporating the concept of momentum, which dynamically estimate both the learning rates and momentum as the number of iterations increases. In our method, the growth of the gradients is controlled using a scaling (clipping) function, leading to stable convergence. Within this framework, we propose three distinct estimators for learning rates and provide theoretical proof of their convergence. We further demonstrate how these estimators underpin the development of effective Nlar optimizers. The performance of the proposed estimators and optimizers is rigorously evaluated through extensive experiments across several datasets and a reinforcement learning environment. The results highlight two key features of the Nlar optimizers: robust convergence despite variations in underlying parameters, including large initial learning rates, and strong adaptability with rapid convergence during the initial epochs.

优化器自适应学习率动态调节

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。