融合动量与随机线搜索,加速大规模深度学习优化
Effectively Leveraging Momentum Terms in Stochastic Line Search Frameworks for Fast Optimization of Finite-Sum Problems
- 用小批量持续性解决动量与线搜索的兼容难题
- 在凸与非凸问题上均达当前最优训练效果
- 适合追求高效训练的大规模模型研究者
本文研究无约束有限求和优化问题,尤其关注大规模深度学习场景。重点探讨过参数化环境下随机优化中的线搜索方法与动量方向的关系。指出二者结合存在计算上的挑战,提出基于小批量持续性的解决方案。构建一种算法框架,融合数据持续性、共轭梯度类动量参数设定及随机线搜索。该算法在合理假设下具备收敛性保证,并在实验中显著优于现有主流方法,在凸与非凸的大规模训练任务中均取得当前最优性能。
原文摘要 · Abstract (English)
In this work, we address unconstrained finite-sum optimization problems, with particular focus on instances originating in large scale deep learning scenarios. Our main interest lies in the exploration of the relationship between recent line search approaches for stochastic optimization in the overparametrized regime and momentum directions. First, we point out that combining these two elements with computational benefits is not straightforward. To this aim, we propose a solution based on mini-batch persistency. We then introduce an algorithmic framework that exploits a mix of data persistency, conjugate-gradient type rules for the definition of the momentum parameter and stochastic line searches. The resulting algorithm provably possesses convergence properties under suitable assumptions and is empirically shown to outperform other popular methods from the literature, obtaining state-of-the-art results in both convex and nonconvex large scale training problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。