证明平滑DNN在梯度下降下能达到最优泛化误差
Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent
- 用核方法近似有限宽DNN的训练过程
- 宽度随样本量多项式增长时达最优泛化率
- 首次为平滑激活的全连接DNN提供理论保障
理解过参数化深度神经网络(DNN)的优化动态与统计性能,仍是解释深度学习成功的关键挑战。本文建立量化边界,表明由确定性无限宽神经正切核诱导的再生核希尔伯特空间中的核梯度下降,可近似有限宽、平滑激活函数的DNN在梯度下降(GD)和随机梯度下降(SGD)下的训练过程。近似误差由网络宽度和训练时长决定,且在SGD情况下额外包含随机梯度误差。这一联系为将核方法的学习理论保证转移到深度回归提供了通用机制。作为应用,在一般源条件与有效维数条件下,只要网络宽度随样本量多项式增长,基于GD和SGD训练的DNN均能达到最小最大最优的过剩总体风险率(对数因子内)。据我们所知,这是首个针对标准全连接平滑激活DNN在GD和SGD训练下的此类理论保证。
原文摘要 · Abstract (English)
Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central challenge in understanding the remarkable success of deep learning. We establish quantitative bounds showing that kernel gradient descent in the reproducing kernel Hilbert space induced by the deterministic infinite-width neural tangent kernel approximates finite-width deep regression with smooth activations under gradient descent (GD) and stochastic gradient descent (SGD) training. The approximation gap is governed by the network width and training horizon, with an additional stochastic gradient error in the SGD case. This connection provides a general mechanism for transferring learning-theoretic guarantees from kernel methods to deep regression. As an application, under general source and effective dimension conditions, we show that both GD- and SGD-trained DNNs attain the minimax-optimal excess population risk rate, up to logarithmic factors, provided that the network width grows polynomially in the sample size. To the best of our knowledge, these are the first such guarantees for standard fully connected deep neural networks with smooth activations trained by GD and SGD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。