arXiv:2502.13110cs.LG2025-02

提出新参数化方法,让模型训练突破稳定边界仍能有效学习特征。

Feature Learning Beyond the Edge of Stability

  • 采用多项式宽度模式的多层感知机,结合深度梯度缩放。
  • 实现突破稳定边界训练,避免损失爆炸且提升特征质量。
  • 适合研究深度学习训练稳定性与特征学习机制的学者。

我们提出一种同质多层感知机参数化方法,采用多项式隐藏层宽度模式,并在一般监督学习场景下分析其在随机梯度下降与深度梯度缩放下的训练动态。推导出小批量损失前三个泰勒系数的表达式,揭示了尖锐性与特征学习之间的联系,特别提出一种软秩变体来量化隐藏层特征的学习质量。基于该理论,设计了一种梯度缩放方案,配合二次宽度模式,可在不引发损失爆炸或数值错误的情况下实现突破稳定边界的训练,实证显示其提升了特征学习能力并带来了隐式的尖锐性正则化。

原文摘要 · Abstract (English)

We propose a homogeneous multilayer perceptron parameterization with polynomial hidden layer width pattern and analyze its training dynamics under stochastic gradient descent with depthwise gradient scaling in a general supervised learning scenario. We obtain formulas for the first three Taylor coefficients of the minibatch loss during training that illuminate the connection between sharpness and feature learning, providing in particular a soft rank variant that quantifies the quality of learned hidden layer features. Based on our theory, we design a gradient scaling scheme that in tandem with a quadratic width pattern enables training beyond the edge of stability without loss explosions or numerical errors, resulting in improved feature learning and implicit sharpness regularization as demonstrated empirically.

深度学习特征学习训练稳定梯度缩放

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。