证明有限宽度网络梯度下降可实现局部线性收敛,关键在NTK条件与损失函数的PL不等式。
From Sublinear to Linear: Local Convergence in Finite-Width Networks via Locally Polyak-Lojasiewicz Regions
- 通过NTK正定与LQCR区域稳定性,推导出局部PL不等式。
- 在二分类MNIST上,损失呈几何衰减,PL比值有正下界。
- 小步长可维持局部收敛,适合研究训练动态与收敛机制的人。
我们研究有限宽度前馈网络在平方经验损失下的梯度下降局部线性收敛性。已有工作表明,梯度下降可保持在初始化附近的局部拟凸区域(LQCR)内,但仅能保证次线性收敛速率。本文证明:若经验神经切线核(NTK)在初始化时正定,在LQCR上李普希茨稳定,且与LQCR半径兼容,则平方损失满足局部Polyak-Łojasiewicz(PL)不等式,常数μ=λ₀−LΘr(ℛ)>0。结合固定步长迭代被限制在LQCR内的假设,可得该区域内线性收敛。LQCR提供局部化,固定步长包含性作为线性速率定理的假设,而PL不等式源于平方损失下的NTK条件。因此结果为充分局部条件,非必要或唯一机制。实验上,通过分析NTK谱间隙、参数漂移、经验PL比值和次优性衰减验证理论。在二分类MNIST上,NTK保持正定,PL比值有正下界,损失呈现几何衰减;宽度消融实验中,步长为固定时宽度1024的运行脱离局部区域;降低步长使最终漂移从1.870降至0.158,恢复局部诊断指标,并获得研究中最大的经验PL比值下界。在CIFAR-10子集上的卷积网络鲁棒性测试显示,三个种子下PL比值下界均保持正值。
原文摘要 · Abstract (English)
We study local linear convergence of gradient descent for finite-width feedforward networks under the squared empirical loss. Prior work shows that GD can remain confined to a Locally Quasi-Convex Region (LQCR) around initialization, but only gives a sublinear rate. We show that if the empirical Neural Tangent Kernel is positive at initialization, Lipschitz stable on the LQCR, and compatible with the LQCR radius, then the squared loss satisfies a local Polyak-Łojasiewicz inequality with constant $μ= λ_0 - L_Θr(\Rcal) > 0$. Combined with fixed-step iterate containment in the LQCR, imposed as a hypothesis in the linear-rate theorem, this yields linear convergence on the region. The LQCR supplies localization; fixed-step containment is imposed as a hypothesis in the linear-rate theorem; and the PL inequality comes from NTK conditioning under squared loss. The result is therefore a sufficient local condition, not a claim that this mechanism is necessary or unique for fast convergence. Empirically, we probe the theory through NTK spectral gap, parameter drift, empirical PL ratio, and suboptimality decay. On binary MNIST, the NTK remains positive, the PL ratio has a positive lower envelope, and the loss shows geometric decay on the stable regime. In a width ablation, the fixed-step width-$1024$ run leaves the local regime; reducing the step size lowers final drift from $1.870$ to $0.158$, restores the observed local-regime diagnostics, and yields the largest empirical PL-ratio lower envelope observed in the study. A CNN robustness check on a CIFAR-10 subset shows the PL-ratio envelope remains positive across three seeds, with a positive lower envelope across all three seeds on the stable regime.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。