arXiv:2409.07401cs.LGmath.OC2024-09被引 4
连续时间梯度下降在深度网络训练中收敛性获理论突破
Convergence of continuous-time stochastic gradient descent with applications to deep neural networks
- 用连续时间模型分析随机梯度下降的收敛机制
- 证明过参数化神经网络训练中算法可收敛
- 为深度学习优化提供新理论工具,适合研究者参考
我们研究了学习问题中最小化总体期望损失的连续时间随机梯度下降过程。主要结果建立了普遍的充分收敛条件,扩展了Chatterjee(2022)针对(非随机)梯度下降所得结论。我们展示了如何将主结果应用于过参数化神经网络训练情形。
原文摘要 · Abstract (English)
We study a continuous-time approximation of the stochastic gradient descent process for minimizing the population expected loss in learning problems. The main results establish general sufficient conditions for the convergence, extending the results of Chatterjee (2022) established for (nonstochastic) gradient descent. We show how the main result can be applied to the case of overparametrized neural network training.
优化理论神经网络随机梯度
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。