统一模型解释神经网络学习突变与缩放规律的机制。
Neural Quadratic Forms: A Unified Minimal Model for Sudden Learning and Scaling Laws

- 基于对称性推导出通用二次形式,仅由结构矩阵决定
- 初始权重越小,学习突变间隔越长,出现平台期
- 适用于多类架构,可预测训练时间的幂律指数
通过梯度下降训练的神经网络在平滑代价函数下仍表现出分步学习:代价函数长期处于平台期后突然下降。与此同时,训练损失则遵循平滑的幂律。不同微观结构的架构均表现出类似行为,表明存在少数关键集体变量。我们发现,网络层由可互换单元之和构成,其重命名不变性约束了初始权重附近展开的主导形式为二次型 $ r[WW^{ op}A(x)]$,所有架构细节被封装于单一结构矩阵 $A(x)$,该矩阵对各类架构可计算。感知机、注意力层、专家混合与卷积层均可归一化为此模型。训练动态收敛至序参量 $M=WW^{ op}$,当数据矩阵共享特征基时,系统简化为洛特卡-沃尔特拉方程,各模式依次开启;初始权重越小,开启时间间隔越大,平台期是平滑流的奇异极限;当大量模式未分辨时,事件合并成训练时间的幂律,理论预测其指数。数值实验验证了跨训练方法与架构的一致性。
原文摘要 · Abstract (English)
Neural networks trained by gradient descent on a smooth cost function can nevertheless learn in steps: the cost holds on long plateaus and then drops abruptly. Meanwhile, training losses instead follow smooth power laws. Variants of both behaviors occur in architectures with very different microscopic structures, which is the signature of a few relevant collective variables. We show that a symmetry fixes what those variables are: a network layer is a sum over interchangeable units, so relabeling the units leaves it unchanged; given smoothness and the condition that a unit's gradient vanish at the origin, symmetry then enforces a universal leading form for the expansion about the near-zero weights present at the start of training, the quadratic $\Tr[WW^{\top}A(x)]$, in which every architectural detail is confined to a single ``structure matrix" $A(x)$ that we compute for each architecture. Perceptrons, attention layers, mixtures of experts, and convolutions become one model at different $A$. Its training dynamics then close on the ``order parameter" $M=WW^{\top}$ and, whenever the data matrices share an eigenbasis, reduce to a Lotka--Volterra equation whose modes switch on one after another. The smaller the initial weights, the further apart the switch-on times, and the plateaus appear as a singular limit of a smooth flow; when many modes are unresolved the same events merge into a power law in training time whose exponent the theory predicts. We confirm both numerically across training methods and architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。