用随机矩阵理论揭示梯度下降的权重演化规律,给出学习率与批量大小的线性关系。
Random Matrix Theory for Stochastic Gradient Descent
- 基于狄森布朗运动框架,用随机矩阵理论分析权重动态
- 推导出学习率与批量大小的线性比例关系
- 在玻尔兹曼机和单层神经网络中验证,适用于深度学习调参
研究机器学习算法的学习动态对理解其成功机制至关重要。物理与统计工具为此类研究提供了坚实框架。本文应用随机矩阵理论描述随机权重矩阵的动力学,采用狄森布朗运动框架。我们推导出学习率(步长)与批量大小之间的线性缩放规则,并识别了权重矩阵动态中的普适与非普适特性。我们的结论在高斯受限玻尔兹曼机(Gaussian Restricted Boltzmann Machine)这一近可解模型及单隐层线性神经网络中得到验证。
原文摘要 · Abstract (English)
Investigating the dynamics of learning in machine learning algorithms is of paramount importance for understanding how and why an approach may be successful. The tools of physics and statistics provide a robust setting for such investigations. Here we apply concepts from random matrix theory to describe stochastic weight matrix dynamics, using the framework of Dyson Brownian motion. We derive the linear scaling rule between the learning rate (step size) and the batch size, and identify universal and non-universal aspects of weight matrix dynamics. We test our findings in the (near-)solvable case of the Gaussian Restricted Boltzmann Machine and in a linear one-hidden-layer neural network.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。