证明了宽神经网络在多项式宽度下可逼近无限宽模型的训练动态。
Propagation of Chaos in One-hidden-layer Neural Networks beyond Logarithmic Time
- 通过描述粒子速度与位置关系的微分方程,严格控制有限宽与无限宽网络的差距。
- 在高维单指标模型中,多项式数量神经元即可在训练全程保持近似性。
- 适用于需要长时间训练的高维学习任务,尤其适合研究理论收敛机制的人。
我们研究在均场标度下,使用投影梯度下降训练的多项式宽度神经网络与其无限宽度版本之间的近似差距。通过一个由均场动力学决定的微分方程,我们严格界定了该差距。影响该常微分方程增长的关键因素是每个粒子的局部海森矩阵,即粒子速度对位置导数。我们将结果应用于经典特征学习问题——估计一个设定良好的单指标模型;允许信息指数任意大,导致收敛时间在环境维度 $d$ 上多项式增长。我们证明,由于此类问题存在某种“自协调”性质——即粒子的局部海森矩阵被其速度的常数倍所控制——因此多项式数量的神经元足以在整个训练过程中紧密逼近均场动力学。
原文摘要 · Abstract (English)
We study the approximation gap between the dynamics of a polynomial-width neural network and its infinite-width counterpart, both trained using projected gradient descent in the mean-field scaling regime. We demonstrate how to tightly bound this approximation gap through a differential equation governed by the mean-field dynamics. A key factor influencing the growth of this ODE is the local Hessian of each particle, defined as the derivative of the particle's velocity in the mean-field dynamics with respect to its position. We apply our results to the canonical feature learning problem of estimating a well-specified single-index model; we permit the information exponent to be arbitrarily large, leading to convergence times that grow polynomially in the ambient dimension $d$. We show that, due to a certain ``self-concordance'' property in these problems -- where the local Hessian of a particle is bounded by a constant times the particle's velocity -- polynomially many neurons are sufficient to closely approximate the mean-field dynamics throughout training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。