证明了浅层网络在无噪声下随时间均匀收敛,突破传统依赖凸性的限制。
Uniform-in-Time Weak Propagation-of-Chaos in Shallow Neural Networks

- 基于均场动力学的梯度流收敛速率,推导出非渐近弱传播混沌界。
- 当均场损失以超 $t^{-2}$ 速率下降时,只需 $\text{poly}(d/ε)$ 个神经元即达精度 $ε$。
- 无需假设优化曲面几何结构,适用于有限样本与离散时间情形。
研究使用梯度下降训练的一层神经网络在特征学习阶段的表现,将有限宽度网络 $f_{\hatρ_t^m}$ 的输出与无限宽度对应物 $f_{ρ_t^{MF}}$(遵循均场动力学)进行比较。尽管通过标准 Grönwall 估计可获得固定时间范围内的误差界,但长期波动行为更复杂。本文在无噪声设置下,通过利用均场确定性 Wasserstein 梯度流的动力学收敛速率,建立了非渐近的、关于时间均匀的弱传播混沌界。在标准正则性假设和 $\int_0^\infty L_t^{1/2} dt = O(\log d)$ 条件下,若 $L_t \lesssim t^{-c}$,则有 $\|f_{ρ_t^{MF}} - f_{\hatρ_t^m}\|^2 \lesssim \text{poly}(d) m^{-\min(1,c/6)}$,其中 $m$ 为神经元数量。该结果不依赖最优解附近的局部强凸性或对数 Sobolev 不等式,且可自然推广至有限样本与时间离散化。核心结论是:当均场种群损失的收敛速率快于 $t^{-2}$ 时,仅需 $\text{poly}(d/ε)$ 个神经元、训练样本和梯度下降步数即可达到精度 $ε$。
原文摘要 · Abstract (English)
We consider one-hidden layer neural networks trained in the feature-learning regime using gradient descent, and relate the output of the finite-width network $f_{\hatρ_t^m}$ to its infinite-width counterpart $f_{ρ_t^{MF}}$, which evolves in the mean-field dynamics. While constant-time horizon bounds for $\|f_{ρ_t^{MF}} - f_{\hatρ_t^m}\|$ may be obtained via standard Grönwall estimates, the long-time behavior of the fluctuation is a more delicate matter. Uniform-in-time bounds often rely on (local) strong convexity in the landscape or Logarithmic Sobolev inequalities present in noisy gradient dynamics. In this work, we establish non-asymptotic weak propagation-of-chaos that holds uniformly in time, obtained by exploiting instead the convergence rate of the mean-field deterministic Wasserstein-gradient-flow dynamics. Specifically, denoting by $L_t$ the mean-field excess MSE loss at time $t$ and $m$ the number of neurons, under standard regularity assumptions and the condition $\int_0^\infty L_t^{1/2} dt =O(\log d)$, we obtain the uniform in time bound $\|f_{ρ_t^{MF}}- f_{\hatρ_t^m}\|^2 \lesssim \text{poly}(d) m^{-\min(1,c/6)}$ whenever $L_t \lesssim t^{-c}$. Our result holds in a noiseless setting and does not make any assumptions on the geometry of the landscape near the optimum, and extends seamlessly to other forms of discretization, including finite number of samples and time discretization. A key takeaway of our result is that whenever the convergence rate of the mean-field, population-loss dynamics is faster than $t^{-2}$, we can attain a loss of $ε$ with only $\text{poly}(d/ε)$ neurons, training samples, and GD steps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。