arXiv:2511.01734stat.MLcs.AI2025-11被引 3

证明了μP参数化下学习率随网络宽度收敛到非零常数

A Proof of Learning Rate Transfer under $μ$P

  • 采用μP参数化构建线性多层感知机,理论推导学习率转移机制
  • 在无限宽极限下,最优学习率收敛至非零常数,解释学习率转移现象
  • 对比传统参数化失效,适合研究神经网络训练理论的读者

我们首次在μP参数化的线性多层感知机(MLP)中证明了学习率转移现象。μP是一种旨在最大化无限宽极限下特征学习能力的神经网络参数化方法。结果表明,在μP下,随着网络宽度趋于无穷,最优学习率收敛至一个非零常数,为学习率转移提供了理论解释。相比之下,标准参数化(SP)和神经正切参数化(NTP)均不满足此性质。我们给出了直观的理论证明,并通过大量实验验证了这些发现。

原文摘要 · Abstract (English)

We provide the first proof of learning rate transfer with width in a linear multi-layer perceptron (MLP) parametrized with $μ$P, a neural network parameterization designed to ``maximize'' feature learning in the infinite-width limit. We show that under $μP$, the optimal learning rate converges to a \emph{non-zero constant} as width goes to infinity, providing a theoretical explanation to learning rate transfer. In contrast, we show that this property fails to hold under alternative parametrizations such as Standard Parametrization (SP) and Neural Tangent Parametrization (NTP). We provide intuitive proofs and support the theoretical findings with extensive empirical results.

学习率理论分析μP参数化深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。