arXiv:2605.10775math.OCcs.LG2026-05被引 1

证明宽浅网络梯度流会收敛到全局最优解,突破了传统限制。

On the global convergence of gradient flow for wide shallow models beyond homogeneous nonlinearities

  • 基于平均场梯度流,构造参数空间逃逸区证明非全局极小点不稳定。
  • 在隐藏单元数趋于无穷时,若梯度流收敛,则极限必为全局最小。
  • 适用于多头注意力、向量输出等更广泛模型,适合理论研究者。

神经网络训练中一个令人惊讶的现象是:尽管损失函数非凸,梯度下降仍能找到全局最小值。本文研究宽浅网络的全局收敛性,扩展了以往仅限于正一阶齐次非线性(如ReLU)或标量输出有界非线性(如Sigmoid)的结论。我们考虑包含多头注意力层和具有有界或渐近正一阶齐次激活函数及向量输出权重的两层网络。借鉴Chizat & Bach (2018) 的框架,证明在隐藏神经元或注意力头数量趋于无穷时,非全局最小值在平均场梯度流下不稳定的机制——通过构建参数空间中的“逃逸区域”。我们的全局收敛结论为条件性:若平均场梯度流在W2范数下收敛,则其极限必为全局最小。我们重新审视了[CB18]中的有界非线性、标量输出情形,提出适配无界参数域的逃逸区域构造;还针对至多线性增长的非线性(在非退化假设下)及渐近正一阶齐次非线性提出了新构造。最后,证明了在次高斯初始化下平均场训练动态的适定性与稳定性估计。

原文摘要 · Abstract (English)

A surprising phenomenon in the training of neural networks is the ability of gradient descent to find global minimizers of the training loss despite its non-convexity. Following earlier work, we investigate this behavior for wide shallow models. Existing global convergence results primarily concern models with positively one-homogeneous nonlinearities, such as ReLU activations, and models with scalar output weights and bounded nonlinearities, such as sigmoid activations. We study a broader class of models, including multi-head attention layers and two-layer networks with bounded or asymptotically positively one-homogeneous activations and vector output weights. Building upon [Chizat and Bach, 2018], we prove that, in the limit of many hidden neurons or attention heads, non-global minimizers of the training loss are unstable under mean-field gradient flow dynamics by constructing "escape regions" in the parameter space. Our global convergence statements are conditional in the following sense: if the mean-field gradient flow converges in W2, then its limit must be a global minimizer. We revisit the bounded nonlinearity, scalar-output setting of [CB18], giving an escape region construction adapted to unbounded nonlinear parameter domains. We also propose new constructions for nonlinearities with at most linear growth under a non-degeneracy assumption and for asymptotically positively one-homogeneous nonlinearities. Finally, we show the well-posedness and stability estimates for the mean-field training dynamics under sub-Gaussian initializations.

深度学习理论梯度流全局收敛

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。