基于终端吸引子设计四种自适应学习率,加速梯度下降收敛。
Dynamic Decoupling of Placid Terminal Attractor-based Gradient Descent Algorithm
- 分阶段利用终端滑模与吸引子理论设计自适应学习率。
- 在函数逼近与图像分类任务中,学习过程总时长显著缩短。
- 适合追求高效优化的机器学习研究者与工程应用开发者。
梯度下降(GD)和随机梯度下降(SGD)在众多应用领域广泛应用,因此深入理解其动态特性并提升收敛速度仍具重要意义。本文针对基于终端吸引子的梯度流在不同阶段的动力学行为进行细致分析。基于终端滑模理论与终端吸引子理论,设计了四种自适应学习率,并通过详尽的理论分析评估其性能,对比了学习过程的运行时间与总迭代次数。为验证有效性,在函数逼近问题与图像分类任务上进行了多组仿真实验,结果表明所提方法在收敛速度与稳定性方面均有显著提升。
原文摘要 · Abstract (English)
Gradient descent (GD) and stochastic gradient descent (SGD) have been widely used in a large number of application domains. Therefore, understanding the dynamics of GD and improving its convergence speed is still of great importance. This paper carefully analyzes the dynamics of GD based on the terminal attractor at different stages of its gradient flow. On the basis of the terminal sliding mode theory and the terminal attractor theory, four adaptive learning rates are designed. Their performances are investigated in light of a detailed theoretical investigation, and the running times of the learning procedures are evaluated and compared. The total times of their learning processes are also studied in detail. To evaluate their effectiveness, various simulation results are investigated on a function approximation problem and an image classification problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。