提出一种加速深度网络非凸优化的新共轭梯度法。
Scaled Conjugate Gradient Method for Nonconvex Optimization in Deep Neural Networks
- 改进共轭梯度法,利用随机梯度加速收敛。
- 在图像与文本分类中比现有自适应方法更快降低损失。
- 生成对抗网络训练中达最低FID分数,适合高精度生成任务。
针对深度神经网络中的非凸优化问题,提出一种改进的缩放共轭梯度法,可加速现有基于随机梯度的自适应方法。理论上证明:无论学习率固定或递减,该方法均可收敛至驻点;且在递减学习率下,收敛速度优于经典共轭梯度法。实际应用中,该方法在图像与文本分类任务中比现有自适应方法更快最小化训练损失。在生成对抗网络训练中,其一个版本在所有自适应方法中取得最低的弗雷谢特起始距离(Frechet Inception Distance, FID)得分。
原文摘要 · Abstract (English)
A scaled conjugate gradient method that accelerates existing adaptive methods utilizing stochastic gradients is proposed for solving nonconvex optimization problems with deep neural networks. It is shown theoretically that, whether with constant or diminishing learning rates, the proposed method can obtain a stationary point of the problem. Additionally, its rate of convergence with diminishing learning rates is verified to be superior to that of the conjugate gradient method. The proposed method is shown to minimize training loss functions faster than the existing adaptive methods in practical applications of image and text classification. Furthermore, in the training of generative adversarial networks, one version of the proposed method achieved the lowest Frechet inception distance score among those of the adaptive methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。