量化了浅层神经网络训练中收敛到高斯过程的速度。
Quantitative convergence of trained single layer neural networks to Gaussian processes
- 用有限宽度上界描述网络输出与高斯近似的距离
- 证明距离随网络宽度呈多项式衰减
- 揭示宽度、输入维度和训练动态的影响机制
本文研究了通过梯度下降训练的浅层神经网络在无限宽极限下向其对应高斯过程的定量收敛性。尽管先前工作已在广泛设定下建立了定性收敛,但对训练过程中的有限宽度估计仍十分有限。本文给出了任意训练时间 $t \ge 0$ 下,网络输出与其高斯近似之间二次 Wasserstein 距离的显式上界,证明该距离随网络宽度呈现多项式衰减。结果量化了架构参数(如宽度和输入维数)对收敛的影响,以及训练动态对近似误差的作用。
原文摘要 · Abstract (English)
In this paper, we study the quantitative convergence of shallow neural networks trained via gradient descent to their associated Gaussian processes in the infinite-width limit. While previous work has established qualitative convergence under broad settings, precise, finite-width estimates remain limited, particularly during training. We provide explicit upper bounds on the quadratic Wasserstein distance between the network output and its Gaussian approximation at any training time $t \ge 0$, demonstrating polynomial decay with network width. Our results quantify how architectural parameters, such as width and input dimension, influence convergence, and how training dynamics affect the approximation error.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。