arXiv:2509.23887cs.LGcs.AI2025-09

提出通用神经网络梯度流线性收敛的统一理论,覆盖多种激活函数。

Gradient Flow Convergence Guarantee for General Neural Network Architectures

  • 基于梯度流建立统一收敛定理,适用于分段非零多项式、ReLU、Sigmoid激活。
  • 在无限小步长下实现线性收敛,理论预测与实际训练高度吻合。
  • 为复杂网络优化提供普适性理论支撑,适合研究者参考。

现代深度学习理论的关键挑战在于解释梯度优化方法在大规模复杂神经网络训练中的卓越表现。尽管已有少数特定架构证明了线性收敛性,但尚无统一理论。本文针对任意具有分段非零多项式激活或ReLU、Sigmoid激活的神经网络,提出了连续梯度下降(即梯度流)线性收敛的统一证明。核心贡献是一条通用定理,不仅涵盖此前未知结果的架构,还在更弱假设下整合了已有成果。虽理论聚焦于无穷小步长极限,但其预测与实际梯度下降方法在实践中表现出极佳一致性。

原文摘要 · Abstract (English)

A key challenge in modern deep learning theory is to explain the remarkable success of gradient-based optimization methods when training large-scale, complex deep neural networks. Though linear convergence of such methods has been proved for a handful of specific architectures, a united theory still evades researchers. This article presents a unified proof for linear convergence of continuous gradient descent, also called gradient flow, while training any neural network with piecewise non-zero polynomial activations or ReLU, sigmoid activations. Our primary contribution is a single, general theorem that not only covers architectures for which this result was previously unknown but also consolidates existing results under weaker assumptions. While our focus is theoretical and our results are only exact in the infinitesimal step size limit, we nevertheless find excellent empirical agreement between the predictions of our result and those of the practical step-size gradient descent method.

优化理论梯度流神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。