arXiv:2606.06764stat.MLcs.AI2026-06

首次证明深度神经网络梯度下降可达到最优泛化率。

Optimal Rates for Generalization of Gradient Descent Methods with Deep Neural Networks

  • 在深度ReLU网络下,分析梯度下降与随机梯度下降的泛化性能。
  • 当网络宽度随深度和样本量多项式增长时,达到最小最大最优泛化率。
  • 结果表明深度网络可媲美核方法,理论覆盖深层架构空白。

近年来,关于过参数化神经网络在神经正切核(NTK)框架下梯度下降方法的统计泛化性能已有进展。然而,现有针对回归问题的研究大多局限于浅层网络结构,深度神经网络的理论仍存在显著空白。本文填补这一空白,对使用梯度下降(GD)和随机梯度下降(SGD)训练的深度ReLU网络进行了全面的泛化分析。具体而言,在网络宽度随网络深度和训练样本量多项式增长的前提下,我们首次建立了深度ReLU网络在GD和SGD下的超额总体风险的最小最大最优率。结果表明,当网络足够宽时,梯度下降方法在深度ReLU网络上能达到与核方法相当的最优泛化性能。

原文摘要 · Abstract (English)

Recent progress has been made in understanding the statistical generalization performance of gradient descent methods for overparameterized neural networks within the neural tangent kernel (NTK) regime. However, most of the existing work on regression problems is limited to shallow network architectures, leaving a notable gap in the theory of deep neural networks. This paper addresses this gap by presenting a comprehensive generalization analysis for deep ReLU networks trained using gradient descent (GD) and stochastic gradient descent (SGD). Specifically, we establish the first known minimax-optimal rates of excess population risk for both GD and SGD with deep ReLU networks, under the assumption that the network width scales polynomially with respect to the network depth and training sample size. Our results demonstrate that with sufficient width, gradient descent methods for deep ReLU networks can achieve optimal generalization rates on par with kernel methods.

深度学习泛化理论梯度下降神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。