通过相图分析梯度下降的权重动态,指导超参数选择
Phase diagram and eigenvalue dynamics of stochastic gradient descent in multilayer neural networks
- 用无序系统类比神经网络,构建权重奇异值的相图
- 训练中出现三种动力学相,对应不同收敛行为
- 为学习率与批量大小搭配提供可落地的调参依据
超参数调优是确保机器学习模型收敛的关键步骤。我们提出,通过研究多层神经网络的相图,可获得对随机梯度下降最优超参数选择的直观理解,其中每个相由权重矩阵奇异值的特定演化特征定义。受无序系统启发,我们将均方误差下的多层神经网络损失景观视为特征空间中的无序系统,其中学习到的特征映射为软自旋自由度,权重初始方差代表无序强度,温度由学习率与批次大小的比值决定。随着训练进行,可识别出三个动力学相,其权重矩阵演化性质截然不同。利用基于迪森布朗运动推导的随机梯度下降朗之万方程,我们有效分类了这三种动态区间,为优化器超参数选择提供了实用指导。
原文摘要 · Abstract (English)
Hyperparameter tuning is one of the essential steps to guarantee the convergence of machine learning models. We argue that intuition about the optimal choice of hyperparameters for stochastic gradient descent can be obtained by studying a neural network's phase diagram, in which each phase is characterised by distinctive dynamics of the singular values of weight matrices. Taking inspiration from disordered systems, we start from the observation that the loss landscape of a multilayer neural network with mean squared error can be interpreted as a disordered system in feature space, where the learnt features are mapped to soft spin degrees of freedom, the initial variance of the weight matrices is interpreted as the strength of the disorder, and temperature is given by the ratio of the learning rate and the batch size. As the model is trained, three phases can be identified, in which the dynamics of weight matrices is qualitatively different. Employing a Langevin equation for stochastic gradient descent, previously derived using Dyson Brownian motion, we demonstrate that the three dynamical regimes can be classified effectively, providing practical guidance for the choice of hyperparameters of the optimiser.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。