提出异步自适应优化方法,证明其在非凸函数上收敛速度达O(1/√t)
Stochastic convergence of parallel asynchronous adaptive first-order methods
- 设计异步自适应一阶优化算法,支持动量与不精确归一化
- 在随机设置下,非凸函数收敛速度为O(1/√t)(含对数因子)
- 适用于异构大规模机器学习系统,实测表现优异
引入一类新的异步自适应一阶优化方法,包含多个流行算法的异步变体。还考虑了使用动量和/或不精确归一化的版本。在完全随机设置下分析了该类方法在非凸函数上的收敛性,证明其收敛速度(至多含对数因子)为O(1/√t),前提假设合理。数值实验表明,这类异步自适应算法在异构大规模机器学习系统中具有高度相关性。
原文摘要 · Abstract (English)
A new class of asynchronous adaptive first-order optimization methods is introduced, comprising asynchronous variants of several popular algorithms. Versions of these methods using momentum and/or inexact normalization are also considered. The convergence of methods in the class on non-convex functions is analyzed in a fully stochastic setting, and is shown to be (up to logarithmic factors) of order O(1/sqrt{t}) under reasonable assumptions. Numerical experiments suggest that such asynchronous adaptive algorithms are very relevant in heterogeneous large-scale machine learning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。