arXiv:2602.06208cs.LGcs.AI2026-02

发现平滑激活的MLP训练时权重会聚集在低维子空间,可大幅压缩模型参数。

Emergent Low-Rank Training Dynamics in MLPs with Smooth Activations

  • 分析带平滑激活的两层MLP梯度下降过程,证明权重动态始终集中在不变低维子空间。
  • 实验证明该现象超出理论假设范围,且低秩参数化能媲美全参数模型性能。
  • 适合关注模型压缩、训练机制解释与高效神经网络设计的研究者。

近期实证研究表明,大规模深度神经网络的训练动态发生在低维子空间中。尽管这激发了对低秩训练、压缩与适配的新研究,但非线性网络中此类动态的理论依据仍有限。本文分析了在梯度下降下的多层感知机(MLPs)学习动态,证明权重动态在整个训练过程中始终集中在不变的低维子空间中。理论上,我们精确刻画了具有平滑非线性激活的两层网络的这些不变子空间,揭示了其形成机制。实验上,我们验证了该现象超越理论假设的适用性。基于这些洞见,我们实证表明:若初始化于合适子空间,存在一种低秩MLP参数化方式,可在多种分类任务上达到与全参数模型相当的分类性能。

原文摘要 · Abstract (English)

Recent empirical evidence has demonstrated that the training dynamics of large-scale deep neural networks occur within low-dimensional subspaces. While this has inspired new research into low-rank training, compression, and adaptation, theoretical justification for these dynamics in nonlinear networks remains limited. %compared to deep linear settings. To address this gap, this paper analyzes the learning dynamics of multi-layer perceptrons (MLPs) under gradient descent (GD). We demonstrate that the weight dynamics concentrate within invariant low-dimensional subspaces throughout training. Theoretically, we precisely characterize these invariant subspaces for two-layer networks with smooth nonlinear activations, providing insight into their emergence. Experimentally, we validate that this phenomenon extends beyond our theoretical assumptions. Leveraging these insights, we empirically show there exists a low-rank MLP parameterization that, when initialized within the appropriate subspaces, matches the classification performance of fully-parameterized counterparts on a variety of classification tasks.

低秩训练MLP梯度下降子空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。