arXiv:2605.27097cs.LGstat.ML2026-05

揭示了正交数据下浅层ReLU网络的渐进学习机制与隐式偏差。

Mildly Overparameterized ReLU Networks on Orthogonal Data: Incremental Learning and Implicit Bias

论文配图:Mildly Overparameterized ReLU Networks on Orthogonal Data: Incremental Learning and Implicit Bias
图 1 · 摘自论文原文
  • 研究小初始化下正交数据的两层ReLU网络梯度流动态。
  • 证明网络宽度满足m ≳ log(n)时几乎必然插值训练数据。
  • 发现学习解的ℓ₂范数接近最优插值解,适合关注泛化机制的研究者。

神经网络的成功训练依赖一阶优化方法,但其理论刻画仍不完整,尤其在弱过参数化情形。本文研究从小初始化出发、基于正交训练数据的两层ReLU网络的梯度流动力学。我们证明当初始规模趋于零时,极限流动收敛至鞍点到鞍点的跳跃过程,揭示出每个鞍点处激活一个新神经元的渐进学习现象。该分析恢复了Dana等(2025, arXiv:2502.16977)的结果:当网络宽度m ≳ log(n)时,网络以高概率插值训练数据,其中n为样本数。此外,该渐进过程使我们推导出新的隐式偏差结果:学习得到的插值器的平方ℓ₂-范数按√n缩放,与最小ℓ₂-范数插值器仅差常数因子。更广泛地,本工作首次为ReLU网络提供了渐进学习过程的严格证明,表明弱过参数化网络可收敛至复杂度与最优插值器同阶的插值解。

原文摘要 · Abstract (English)

The successful training of neural networks hinges on the use of first order optimization methods, yet the theoretical characterization of these methods remains incomplete. This is especially true in settings with mild overparameterization. In this work, we study the gradient flow dynamics of two-layer ReLU networks from small initialization with orthogonal training data. We prove the limiting flow converges to a saddle-to-saddle jump process as the initialization scale tends to zero, revealing an incremental learning phenomenon in which a new neuron activates at each saddle. This analysis recovers the known result of Dana et al. (2025, arXiv:2502.16977) that the network interpolates the training data with high probability as soon as $m \gtrsim \log(n)$, where $m$ is the network width and $n$ is the number of training samples. This incremental process characterization also allows us to derive a novel implicit bias result: the learned interpolator has a squared $\ell_2$-norm scaling as $\sqrt{n}$, which is within a constant factor of the minimal $\ell_2$-norm interpolator. More broadly, our work provides the first rigorous proof of an incremental learning process for ReLU networks, whilst suggesting mildly overparameterized networks can converge to interpolating solutions whose complexity is of the same order as that of the optimal interpolator.

深度学习梯度流隐式偏差神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。