arXiv:2606.23364cs.LGmath.DS2026-06

突破传统NTK框架,证明通用神经网络梯度下降收敛性

Convergence of Gradient Descent for General Neural Network Architectures Beyond the NTK Regime

  • 基于网络模块层面构建分析框架,覆盖预归一化多层Transformer等架构
  • 在几乎所有初始化下,梯度下降可收敛至驻点邻域,无需依赖NTK假设
  • 揭示学习率与深度、有效瓶颈维数相关,适用于实际训练场景

训练动态是理解神经网络的核心,但其理论分析即使对简单架构也十分困难,对现代通用架构更是挑战重重。本文提出一个广义神经网络架构与数据集下的梯度下降(GD)收敛分析框架,突破了传统的神经正切核(NTK)范式。该框架在网络模块层面构建,涵盖预归一化多层Transformer等结构。在温和假设下,证明对于几乎所有初始化,使用常规学习率的梯度下降均能收敛至驻点邻域。主要通过解析性与零测集论证建立迭代依赖的PL型不等式,并借助多项式广义光滑性及局部松弛耗散条件证明轨迹上的Lipschitz光滑性。进一步在Xavier初始化与实际架构缩放下解释定理,表明学习率尺度取决于深度和有效瓶颈维数,而非最大宽度。最后推导出残差连接与函数复合的结构非退化性含义,并在框架内给出全局极小值的通用刻画。

原文摘要 · Abstract (English)

Training dynamics is central to understanding neural networks, yet its theoretical analysis remains difficult even for simple architectures and becomes substantially more challenging for general modern architectures. In this paper, we propose a convergence framework for analyzing gradient descent (GD) dynamics under a broad family of neural network architectures and datasets beyond the neural tangent kernel (NTK) regime. The framework is formulated at the level of network blocks and covers architectures including pre-normalized multi-layer transformers. More precisely, under mild assumptions, we prove that for almost all initializations, GD with regular learning rates converges to the neighbourhood of a stationary point. This is mainly proved by establishing an iterate-dependent PL-type inequality through analyticity and measure-zero arguments, and by proving Lipschitz smoothness along the GD trajectory through polynomial generalized smoothness and a local relaxed dissipative condition. We further interpret the theorem under Xavier initialization and practical architectural scaling, showing that the learning rate scale depends on the depth and effective bottleneck dimensions rather than the largest width. Finally, we derive structural nondegeneracy implications for residual connections and function composition, and provide a generic characterization of global minimizers within our framework.

梯度下降神经网络收敛性Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。