arXiv:2411.13512cond-mat.dis-nncs.LG2024-11中稿 · NeurIPS被引 2

用随机矩阵理论揭示训练中权重矩阵的动态规律

Dyson Brownian motion and random matrix dynamics of weight matrices during learning

  • 将权重更新视为迪森布朗运动,揭示特征值排斥现象
  • 学习率与小批量大小比值决定随机性,解释线性缩放规律
  • 在Transformer和受限玻尔兹曼机中验证了该模型的有效性

训练过程中,机器学习架构的权重矩阵通过随机梯度下降或其变体进行更新。本文采用随机矩阵理论分析由此产生的随机矩阵动力学。首先证明该动力学可普遍描述为迪森布朗运动,导致特征值排斥等现象。随机性程度取决于学习率与小批量大小的比值,解释了经验观察到的线性缩放规则。在受限玻尔兹曼机中验证了这一线性缩放关系。随后研究了Transformer(微型GPT)中的权重矩阵动态,发现其特征值分布从初始化时的马尔琴科-帕斯图尔分布,演变为学习结束时与额外结构相结合的形态。

原文摘要 · Abstract (English)

During training, weight matrices in machine learning architectures are updated using stochastic gradient descent or variations thereof. In this contribution we employ concepts of random matrix theory to analyse the resulting stochastic matrix dynamics. We first demonstrate that the dynamics can generically be described using Dyson Brownian motion, leading to e.g. eigenvalue repulsion. The level of stochasticity is shown to depend on the ratio of the learning rate and the mini-batch size, explaining the empirically observed linear scaling rule. We verify this linear scaling in the restricted Boltzmann machine. Subsequently we study weight matrix dynamics in transformers (a nano-GPT), following the evolution from a Marchenko-Pastur distribution for eigenvalues at initialisation to a combination with additional structure at the end of learning.

随机矩阵深度学习权重动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。