arXiv:2505.22491cs.LGcs.AI2025-05NeurIPS被引 6

大学习率下网络仍能稳定训练,因存在可控发散的中间态。

On the Surprising Effectiveness of Large Learning Rates under Standard Width Scaling

  • 发现大学习率下存在可控发散亚区间,梯度与激活保持稳定。
  • 在交叉熵损失下,无限宽极限中隐藏层特征仍可有效演化。
  • 适用于理解标准初始化下的学习率最优选择,尤其适合多模态任务。

缩放极限(如无限宽极限)是研究大规模模型的有力理论工具。然而,普遍认为现有无限宽理论无法准确解释实际网络行为,特别是采用标准参数化(SP,即He初始化与全局学习率)训练的网络。例如,现有理论预测:大学习率下不稳定,小学习率下特征学习消失。但实践中,最优学习率衰减速度慢于理论预测,且网络在极大宽度下仍具稳定训练与显著特征学习能力。本文表明,此差异并非全由有限宽效应导致。通过精细分析此前被认为不稳定的区域,我们发现该区域包含两个子区间:灾难性不稳定和更温和的可控发散。在交叉熵损失下,后者表现为对数输出发散但梯度与激活保持稳定。在可控发散区边缘的大学习率下,存在明确的无限宽极限,所有隐藏层特征持续演化。跨优化器、架构与数据模态的实验验证:神经网络在交叉熵损失下运行于该可控发散区,而在均方误差损失下则不然。实证结果表明,宽度缩放对预测最大稳定学习率指数具有意外实用性,为最优学习率选择提供指导。最后,我们的分析澄清了近期提出的逐层学习率缩放在标准初始化下的有效性与局限。

原文摘要 · Abstract (English)

Scaling limits, such as infinite-width limits, serve as promising theoretical tools to study large-scale models. However, it is widely believed that existing infinite-width theory does not faithfully explain the behavior of practical networks, especially those trained in standard parameterization (SP) meaning He initialization with a global learning rate. For instance, existing theory for SP predicts instability at large learning rates and vanishing feature learning at stable ones. In practice, however, optimal learning rates decay slower than theoretically predicted and networks exhibit both stable training and non-trivial feature learning, even at very large widths. Here, we show that this discrepancy is not fully explained by finite-width phenomena. Instead, we find a resolution through a finer-grained analysis of the regime previously considered unstable and therefore uninteresting. In particular, we show that, under cross-entropy (CE) loss, the unstable regime comprises two distinct sub-regimes: a catastrophically unstable regime and a more benign controlled divergence regime, where logits diverge but gradients and activations remain stable. Moreover, under large learning rates at the edge of the controlled divergence regime, there exists a well-defined infinite width limit where features continue to evolve in all the hidden layers. In experiments across optimizers, architectures, and data modalities, we validate that neural networks operate in this controlled divergence regime under CE loss but not under MSE loss. Our empirical evidence suggests that width-scaling considerations are surprisingly useful for predicting empirically maximal stable learning rate exponents which provide useful guidance on optimal learning rate exponents. Finally, our analysis clarifies the effectiveness and limitations of recently proposed layerwise learning rate scaling for standard initialization.

深度学习学习率无限宽度交叉熵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。