arXiv:2410.20119cs.LG2024-10

揭示神经网络训练中三阶段损失动态机制,解析平台期成因与突破路径。

On Multi-Stage Loss Dynamics in Neural Networks: Mechanisms of Plateau and Descent Stages

  • 通过分析小初始化下的训练过程,识别出三个损失变化阶段。
  • 发现平台期由梯度稀疏和权重分布不均导致,耗时长且难突破。
  • 引入Wasserstein距离追踪参数分布演化,连接全局趋势与局部调整。

神经网络训练中的多阶段损失现象已被广泛观察到,反映了训练过程的非线性与复杂性。本文聚焦小初始化情形,识别出训练过程中损失曲线的三个显著阶段:初始平台期、初始下降期与二次平台期。通过严格分析,揭示了平台期训练缓慢的深层原因。尽管初始平台期的出现已有理论证明,但初始下降期与二次平台期的行为此前未被系统研究。本文对初始平台期提供更详尽证明,并深入分析初始下降期的动力学特性。进一步,结合实验与启发式推理,探讨了网络克服长期二次平台的关键因素。最后,为厘清全局训练趋势与局部参数更新之间的关联,采用Wasserstein距离追踪权重幅度分布的细粒度演化过程。

原文摘要 · Abstract (English)

The multi-stage phenomenon in the training loss curves of neural networks has been widely observed, reflecting the non-linearity and complexity inherent in the training process. In this work, we investigate the training dynamics of neural networks (NNs), with particular emphasis on the small initialization regime, identifying three distinct stages observed in the loss curve during training: the initial plateau stage, the initial descent stage, and the secondary plateau stage. Through rigorous analysis, we reveal the underlying challenges contributing to slow training during the plateau stages. While the proof and estimate for the emergence of the initial plateau were established in our previous work, the behaviors of the initial descent and secondary plateau stages had not been explored before. Here, we provide a more detailed proof for the initial plateau, followed by a comprehensive analysis of the initial descent stage dynamics. Furthermore, we examine the factors facilitating the network's ability to overcome the prolonged secondary plateau, supported by both experimental evidence and heuristic reasoning. Finally, to clarify the link between global training trends and local parameter adjustments, we use the Wasserstein distance to track the fine-scale evolution of weight amplitude distribution.

训练动态损失曲线神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。