arXiv:2511.09810cs.LGcs.AI2025-11

揭示过参数神经网络优化本质:结构决定收敛性与速度。

On the Convergence of Overparameterized Problems: Inherent Properties of the Compositional Structure of Neural Networks

  • 分析线性激活下梯度流,证明任意合理损失函数全局收敛。
  • 过参数表示决定鞍点位置与稳定性,与具体任务无关。
  • 初始化不平衡可加速收敛,为调参提供新思路。

本文研究神经网络的组合结构如何影响其优化景观与训练动态。我们分析了过参数化优化问题对应的梯度流,可视为训练具有线性激活的神经网络。显著地,我们证明了对于任意合适的实解析损失函数,全局收敛性质均可推导得出。随后将分析聚焦于标量损失函数情形,此时景观几何可被完全刻画。在此设定下,我们证明关键结构性特征——如鞍点的位置与稳定性——对所有允许的损失函数是普遍的,仅依赖于过参数化表示,而非具体问题细节。此外,我们展示了收敛速度可通过本文提出的失衡度量随初始化任意加速。最后,我们讨论这些见解如何推广至带Sigmoid激活的神经网络,并通过一个简单示例表明,某些几何与动力学特性在非线性情况下依然存在。

原文摘要 · Abstract (English)

This paper investigates how the compositional structure of neural networks shapes their optimization landscape and training dynamics. We analyze the gradient flow associated with overparameterized optimization problems, which can be interpreted as training a neural network with linear activations. Remarkably, we show that the global convergence properties can be derived for any cost function that is proper and real analytic. We then specialize the analysis to scalar-valued cost functions, where the geometry of the landscape can be fully characterized. In this setting, we demonstrate that key structural features -- such as the location and stability of saddle points -- are universal across all admissible costs, depending solely on the overparameterized representation rather than on problem-specific details. Moreover, we show that convergence can be arbitrarily accelerated depending on the initialization, as measured by an imbalance metric introduced in this work. Finally, we discuss how these insights may generalize to neural networks with sigmoidal activations, showing through a simple example which geometric and dynamical properties persist beyond the linear case.

优化理论过参数化神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。