发现语言模型预训练中奇异值分布稳定现象,解释了训练为何先快后慢。
The Stability of Singular Distribution: A Spectral Perspective on the Two-Phase Dynamics of Language Model Pre-training

- 通过奇异值谱分析,揭示训练早期奇异值分布即趋于稳定。
- 该稳定现象与缓慢下降阶段同步,在多种架构和优化器下普遍成立。
- 为WSD、Muon等高效训练策略提供了新的谱视角解释。
大型语言模型预训练通常呈现两阶段轨迹:初期快速损失下降,随后进入漫长缓慢提升。我们识别出一种潜在的谱现象——奇异值分布稳定性(SoSD),即迹归一化的奇异值谱在早期即趋于稳定,尽管参数矩阵仍在持续演化。我们证明,SoSD与慢速下降阶段的同步在多种架构(GPT-2、LLaMA)及不同设置(分段、WSD、余弦衰减)下广泛存在,涵盖不同权重衰减和优化器(AdamW、Muon)。通过对简化Transformer的分析,我们证明权重范数增长会不可避免地触发早期的SoSD阈值,此后损失下降速率理论上受限于奇异值分布的变化幅度。我们进一步通过调节SoSD尺度的能力,阐释了WSD与Muon等策略的作用机制,为理解高效预训练动态提供了谱视角。
原文摘要 · Abstract (English)
Large language model pre-training typically exhibits a two-phase trajectory: a fast initial loss drop followed by a prolonged slow improvement. We identify an underlying spectral phenomenon, Stability of Singular Distribution (SoSD), where the trace-normalized singular value spectrum stabilizes early, even as parameter matrices continue to evolve. We demonstrate that synchronization between SoSD and the slow-descent regime is widely observed across diverse architectures (GPT-2, LLaMA) and settings, including various schedules (Step-wise, WSD, Cosine Decay), weight decays, and optimizers (AdamW, Muon). By analyzing a simplified Transformer, we prove that growing weight norms inevitably precipitate an early SoSD threshold, after which the rate of loss decrease becomes theoretically bounded by the variation in the singular distribution. We further interpret strategies like WSD and Muon through their ability to modulate the SoSD scale, offering a spectral lens for understanding efficient pre-training dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。