arXiv:2505.09572cs.LGmath.LO2025-05NeurIPS被引 5

证明神经网络梯度下降会发散或收敛到最优值,且初始接近最优即能逼近零损失。

SAD Neural Networks: Divergent Gradient Flows and Asymptotic Optimality via o-minimal Structures

  • 利用o-极小结构几何分析梯度流行为
  • 初始化足够好时损失趋近零,但仅能渐进达到
  • 理论结合实验验证发散与渐近最优性

我们研究了使用常见连续可微激活函数(如logistic、tanh、softplus或GELU)的全连接前馈神经网络的损失曲面梯度流。证明梯度流要么收敛至临界点,要么发散至无穷,同时损失趋于一个渐近临界值。进一步证明存在阈值ε>0,使得在最优值之上不超过ε处初始化的梯度流,其损失值将收敛至最优值。对于多项式目标函数,在足够大的模型规模和数据集下,最优损失值为零,且只能渐进实现。由此得出主结论:初始化足够优的梯度流将发散至无穷。证明依赖于o-极小结构的几何性质。数值实验验证了这些理论结果,并扩展至更现实场景,观察到类似行为。

原文摘要 · Abstract (English)

We study gradient flows for loss landscapes of fully connected feedforward neural networks with commonly used continuously differentiable activation functions such as the logistic, hyperbolic tangent, softplus or GELU function. We prove that the gradient flow either converges to a critical point or diverges to infinity while the loss converges to an asymptotic critical value. Moreover, we prove the existence of a threshold $\varepsilon>0$ such that the loss value of any gradient flow initialized at most $\varepsilon$ above the optimal level converges to it. For polynomial target functions and sufficiently big architecture and data set, we prove that the optimal loss value is zero and can only be realized asymptotically. From this setting, we deduce our main result that any gradient flow with sufficiently good initialization diverges to infinity. Our proof heavily relies on the geometry of o-minimal structures. We confirm these theoretical findings with numerical experiments and extend our investigation to more realistic scenarios, where we observe an analogous behavior.

神经网络梯度流优化理论o-极小

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。