不同随机种子下语言模型的收敛性差异揭示了模型稳定性关键因素
Convergence and Divergence of Language Models under Different Random Seeds
- 用每词KL散度衡量多种子训练下的收敛性变化
- 发现四阶段演化模式,大模型后期更快重收敛
- 高频词和虚词比低频词和实词收敛更稳定
本文研究了在不同随机种子下语言模型(LMs)的收敛性,将收敛性定义为跨种子的每词平均Kullback-Leibler(KL)散度。通过对比模型规模与训练检查点对收敛性的影响,识别出四个阶段的收敛模式:(i) 初期均一阶段,(ii) 快速收敛阶段,(iii) 快速发散阶段,(iv) 缓慢重收敛阶段。进一步发现,大模型在后期训练中能更快重收敛,而小模型则始终无法真正重收敛;这表明学习稳定分布需要一定模型规模。若限定分析特定词频或词性(PoS)标签,则发现收敛性在语言类别间不均等:高频词与功能词比低频词与内容词收敛更快且更可靠。总体而言,本研究揭示了影响模型训练中学习分布稳定性的关键因素。
原文摘要 · Abstract (English)
In this paper, we investigate the convergence of language models (LMs) trained under different random seeds, measuring convergence as the expected per-token Kullback--Leibler (KL) divergence across seeds. By comparing LM convergence as a function of model size and training checkpoint, we identify a four-phase convergence pattern: (i) an initial uniform phase, (ii) a sharp-convergence phase, (iii) a sharp-divergence phase, and (iv) a slow-reconvergence phase. Further, we observe that larger models reconverge faster in later training stages, while smaller models never actually reconverge; these results suggest that a certain model size may be necessary to learn stable distributions. Restricting our analysis to specific token frequencies or part-of-speech (PoS) tags further reveals that convergence is uneven across linguistic categories: frequent tokens and function words converge faster and more reliably than their counterparts (infrequent tokens and content words). Overall, our findings highlight factors that influence the stability of the learned distributions in model training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。