提出双时间尺度学习法,加速两层神经网络特征学习并保证收敛。
Ultra-fast feature learning for the training of two-layer neural networks in the two-timescale regime
- 采用变量投影法分离线性与非线性参数,简化训练过程。
- 在教师-学生场景下,特征分布收敛速度达到量化保证。
- 适合研究神经网络动态与优化机制的学者参考。
我们研究了均值场单隐层神经网络在平方损失下的梯度方法收敛性。针对这一高维非凸优化问题,现有多数结果为定性或依赖于神经正切核分析(其中数据的非线性表示固定)。由于该问题属于可分离非线性最小二乘问题,本文采用变量投影(VarPro)或双时间尺度学习算法,消除线性变量,将学习问题转化为非线性特征训练。在教师-学生设定下,证明该策略可实现对教师特征分布采样的可证明收敛速率。当正则化强度趋于零时,特征分布的动态演化对应加权超快扩散方程。近期关于此类偏微分方程渐近行为的结果,给出了所学特征分布收敛的定量保证。
原文摘要 · Abstract (English)
We study the convergence of gradient methods for the training of mean-field single-hidden-layer neural networks with square loss. For this high-dimensional and non-convex optimization problem, most known convergence results are either qualitative or rely on a neural tangent kernel analysis where nonlinear representations of the data are fixed. Using that this problem belongs to the class of separable nonlinear least squares problems, we consider here a Variable Projection (VarPro) or two-timescale learning algorithm, thereby eliminating the linear variables and reducing the learning problem to the training of nonlinear features. In a teacher-student scenario, we show such a strategy enables provable convergence rates for the sampling of a teacher feature distribution. Precisely, in the limit where the regularization strength vanishes, we show that the dynamic of the feature distribution corresponds to a weighted ultra-fast diffusion equation. Recent results on the asymptotic behavior of such PDEs then give quantitative guarantees for the convergence of the learned feature distribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。