Adam和SGD训练深度网络时无法收敛到最优风险值
Non-convergence to the optimal risk for Adam and stochastic gradient descent optimization in the training of deep neural networks
- 证明多种SGD类优化器在全连接网络中无法收敛到最优风险
- 真实风险可能收敛至严格次优值而非全局最优
- 适用于研究深度学习优化理论的学者与工程师
尽管随机梯度下降(SGD)在深度神经网络(DNN)训练中被广泛使用,但其成功与局限性的严格理论解释仍是基本开放问题。特别是,尚未能证明或证伪各类优化器的真实风险是否收敛至最优真实风险值。本文主要结果表明:对于一类通用激活函数、损失函数、随机初始化及多种常见优化器(包括标准SGD、动量SGD、Nesterov加速SGD、Adagrad、RMSprop、Adadelta、Adam、Adamax、Nadam、Nadamax、AMSGrad),在任意全连接前馈DNN的训练中,所考虑优化器的真实风险均不以概率收敛至最优真实风险值。尽管如此,真实风险仍可能收敛于一个严格次优的值。
原文摘要 · Abstract (English)
Despite the omnipresent use of stochastic gradient descent (SGD) optimization methods in the training of deep neural networks (DNNs), it remains, in basically all practically relevant scenarios, a fundamental open problem to provide a rigorous theoretical explanation for the success (and the limitations) of SGD optimization methods in deep learning. In particular, it remains an open question to prove or disprove convergence of the true risk of SGD optimization methods to the optimal true risk value in the training of DNNs. In one of the main results of this work we reveal for a general class of activations, loss functions, random initializations, and SGD optimization methods (including, for example, standard SGD, momentum SGD, Nesterov accelerated SGD, Adagrad, RMSprop, Adadelta, Adam, Adamax, Nadam, Nadamax, and AMSGrad) that in the training of any arbitrary fully-connected feedforward DNN it does not hold that the true risk of the considered optimizer converges in probability to the optimal true risk value. Nonetheless, the true risk of the considered SGD optimization method may very well converge to a strictly suboptimal true risk value.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。