用线性模型解析神经网络训练轨迹的低维结构成因
An Analytical Characterization of Sloppiness in Neural Networks: Insights from Linear Models
- 基于动力系统理论分析线性网络训练轨迹的几何特性
- 发现输入相关矩阵特征值衰减速率等三因素决定低维流形存在
- 结果对理解深度网络训练规律有启发,适合研究优化机制者阅读
近期实验表明,不同架构、优化算法、超参数和正则化方法的深度神经网络在概率分布空间中的训练轨迹会演化为显著低维的‘超带状’流形。受深度网络与线性网络训练轨迹相似性的启发,本文针对后者进行解析性刻画。利用动力系统理论,我们证明该低维流形的几何特性由三个因素控制:(i) 训练数据输入相关矩阵特征值的衰减速率,(ii) 初期真值输出与权重的相对尺度,(iii) 梯度下降步数。通过解析计算并界定这些量的贡献,我们刻画了超带状出现的相变边界。分析还扩展至核机器和随机梯度下降训练的线性模型。
原文摘要 · Abstract (English)
Recent experiments have shown that training trajectories of multiple deep neural networks with different architectures, optimization algorithms, hyper-parameter settings, and regularization methods evolve on a remarkably low-dimensional "hyper-ribbon-like" manifold in the space of probability distributions. Inspired by the similarities in the training trajectories of deep networks and linear networks, we analytically characterize this phenomenon for the latter. We show, using tools in dynamical systems theory, that the geometry of this low-dimensional manifold is controlled by (i) the decay rate of the eigenvalues of the input correlation matrix of the training data, (ii) the relative scale of the ground-truth output to the weights at the beginning of training, and (iii) the number of steps of gradient descent. By analytically computing and bounding the contributions of these quantities, we characterize phase boundaries of the region where hyper-ribbons are to be expected. We also extend our analysis to kernel machines and linear models that are trained with stochastic gradient descent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。