证明了深度ReLU网络梯度下降可达到近最优泛化率。
Optimal Rates for Generalization of Gradient Descent for Deep ReLU Classification
- 通过权衡优化与泛化误差,设计新分析方法
- 在数据满足NTK间隔条件下,实现$ ilde{O}(L^6/(nγ^2))$风险率
- 适用于研究深度网络泛化理论的学者
近期研究显著提升了对深度神经网络中梯度下降(GD)泛化性能的理解。一个核心问题是:GD能否在深度网络中达到核方法下建立的极小极大最优泛化率?现有结果或仅得次优率$O(1/ ext{√}n)$,或局限于光滑激活函数,导致深度$L$上呈指数依赖。本文在数据为NTK可分且间隔为$γ$的假设下,通过精细权衡优化与泛化误差,建立了深度ReLU网络梯度下降的最优泛化率,仅含深度的多项式依赖。具体而言,我们证明了超出风险率为$ ilde{O}(L^6 / (n γ^2))$,与最优的SVM型率$ ilde{O}(1 / (n γ^2))$仅差深度相关因子。关键技术贡献在于对参考模型附近激活模式的新控制,从而获得深度ReLU网络训练时更紧的Rademacher复杂度界。
原文摘要 · Abstract (English)
Recent advances have significantly improved our understanding of the generalization performance of gradient descent (GD) methods in deep neural networks. A natural and fundamental question is whether GD can achieve generalization rates comparable to the minimax optimal rates established in the kernel setting. Existing results either yield suboptimal rates of $O(1/\sqrt{n})$, or focus on networks with smooth activation functions, incurring exponential dependence on network depth $L$. In this work, we establish optimal generalization rates for GD with deep ReLU networks by carefully trading off optimization and generalization errors, achieving only polynomial dependence on depth. Specifically, under the assumption that the data are NTK separable from the margin $γ$, we prove an excess risk rate of $\widetilde{O}(L^6 / (n γ^2))$, which aligns with the optimal SVM-type rate $\widetilde{O}(1 / (n γ^2))$ up to depth-dependent factors. A key technical contribution is our novel control of activation patterns near a reference model, enabling a sharper Rademacher complexity bound for deep ReLU networks trained with gradient descent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。