简单二次模型竟可精准预测大模型训练动态
A Defense of the Quadratic Model

- 用中间检查点展开模型与损失函数,预测长达10%训练期的优化过程
- 发现海森矩阵谱在尾部存在结构,受批大小和预处理器影响
- 实证显示大模型训练处于随机稳定边缘,适合理论分析
由于神经网络损失曲面的复杂性,优化理论常依赖理想化模型,但理论可处理性与真实动态描述精度间存在权衡。本文对最简优化模型——二次模型进行压力测试,在包含150M参数和3B训练标记的大语言模型设置中,展示其出人意料的预测能力。具体而言,通过在训练中途检查点对模型和损失函数进行泰勒展开,可准确预测长达训练周期10%的优化动态。在此基础上,我们从海森矩阵谱和局部稳定性两方面分析这些局部二次优化问题。利用极深探针的Lanczos求积法,成功估算出海森矩阵谱的深层尾部,发现特征值与特征向量均呈现显著结构,且受批大小、预处理器及训练时间影响。此外,我们在中间检查点实证检验局部线性稳定性,并与理论预测对比,证实大模型训练通常发生在随机稳定边缘,其性质亦由批大小决定。结果表明,二次模型可能是预训练优化动态的一个理论可处理代理。
原文摘要 · Abstract (English)
Due to the complexity of neural network loss landscapes, optimization theory is forced to rely on idealized models, and there is generally a tradeoff between how theoretically tractable the model is, and how accurately it describes the true optimization dynamics. In this work, we stress test the simplest possible model of optimization -- the quadratic model -- and show that it can be surprisingly predictive in an LLM setting with 150M parameters and 3B training tokens. Specifically, we show that Taylor expanding the model and the loss function at intermediate checkpoints through training can accurately predict the optimization dynamics over windows that can last up to 10\% of training. Having established this agreement, we then turn to analyzing the structure of these local quadratic optimization problems through two lenses: the Hessian spectrum and local stability. Using Lanczos quadrature with extremely deep probes, we are able to estimate the Hessian spectrum deep into the tail, and we find a surprising amount of structure in both the eigenvalues and eigenvectors, which depends on the batch size, preconditioner, and training time. We also empirically test local linear stability at intermediate checkpoints and compare it to theoretical predictions to demonstrate that optimization in LLMs typically occurs at a stochastic edge of stability, whose nature is also determined by batch size. Our results indicate the quadratic model may be a theoretically tractable proxy for pretraining optimization dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。