提出三因素缩放定律,明确训练步数与批大小对模型性能的影响。
How to Allocate Your Tokens? Scaling Laws with Training Steps and Batch Size

- 将训练数据拆分为训练步数与批大小,建立三因素缩放模型。
- 实验验证最优批大小的缩放规律,且仅需少量训练运行即可拟合。
- 可推导次优批大小的缩放关系,与已有实证发现一致。
我们提出一种缩放定律,同时考虑模型规模和训练数据量,并显式区分训练数据为训练步数与批大小(称为三因素定律)。在大量训练运行上拟合该定律,结果正确恢复了最优批大小的缩放规律。由于利用了次优批大小的训练运行,所提定律可用更少的训练运行稳健拟合。此外,三因素定律可用于推导次优批大小的缩放规律,其结果与先前关于临界批大小的实证发现相符。
原文摘要 · Abstract (English)
We propose a scaling law that takes into account model size and training data while explicitly splitting the latter into training steps and batch size (called three-term law). Fitting the proposed law on a large set of training runs, we find that it correctly recovers the scaling of the optimal batch size. Moreover, because it makes use of training runs with suboptimal batch size, our proposed law can be robustly fit with a significantly smaller amount of training runs. We further show that the three-term law can be used to derive scaling laws for suboptimal batch sizes, and that it matches previous empirical findings related to the critical batch size.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。