证明了非线性神经网络梯度下降收敛极慢,仅约1/ln(t)。
On the Rate of Convergence of GD in Non-linear Neural Networks: An Adversarial Robustness Perspective
- 分析双神经元ReLU网络的梯度轨迹,揭示收敛机制。
- 收敛速率严格为Θ(1/ln(t)),远低于实用要求。
- 结果对多种初始化均成立,适用于关注鲁棒性的研究者。
我们在一个最小化的二分类设定中研究梯度下降(GD)的收敛动态,模型由两个神经元的ReLU网络和两个训练样本构成。我们证明,在这些强简化假设下,尽管GD能成功收敛到最优鲁棒性边界,即有效最大化决策边界与训练点之间的距离,但其收敛速率极为缓慢,严格为Θ(1/ln(t))。据我们所知,这是首个在非线性模型中对鲁棒性边界收敛速率的显式下界。通过实验模拟,我们进一步验证这一内在缺陷普遍存在,在多个自然网络初始化下均表现出相同的紧密收敛速率。理论推导基于对模型不同激活模式下GD轨迹的严格分析,通过精细控制系统的动态行为,对决策边界的轨迹进行上界约束,克服了非线性架构带来的主要技术挑战。
原文摘要 · Abstract (English)
We study the convergence dynamics of Gradient Descent (GD) in a minimal binary classification setting, consisting of a two-neuron ReLU network and two training instances. We prove that even under these strong simplifying assumptions, while GD successfully converges to an optimal robustness margin, effectively maximizing the distance between the decision boundary and the training points, this convergence occurs at a prohibitively slow rate, scaling strictly as $Θ(1/\ln(t))$. To the best of our knowledge, this establishes the first explicit lower bound on the convergence rate of the robustness margin in a non-linear model. Through empirical simulations, we further demonstrate that this inherent failure mode is pervasive, exhibiting the exact same tight convergence rate across multiple natural network initializations. Our theoretical guarantees are derived via a rigorous analysis of the GD trajectories across the distinct activation patterns of the model. Specifically, we develop tight control over the system's dynamics to bound the trajectory of the decision boundary, overcoming the primary technical challenge introduced by the non-linear nature of the architecture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。