小模型也能验证缩放定律,关键在调好超参数。
Small-Scale Experiments: Are We There Yet?
- 发现小模型敏感源于超参数,规模增大后影响减弱。
- 调优后的小模型可复现大模型的缩放规律,证明其有效性。
- 适合想用小成本验证大模型结论的研究者参考。
缩放定律曾承诺低成本实验,但六年过去仍难实现。研究发现,在400万参数以上的小规模模型中,缩放定律不可靠,且被认为需大模型才能成立。我们指出问题根源在于超参数:小模型对超参数极为敏感,但这种敏感性随规模增长而减弱。缩放定律仅在完全调优的前沿出现,而达到该前沿需远超常规的搜索范围。通过消融分析,我们证明良好调优的超参数比任何其他因素更重要。进一步揭示,随着规模增加,超参数损失面维度降低,使最优解更易找到。尽管小模型中存在缩放定律,但外推受统计限制。综合新见解与近期文献,我们提出以模型为中心的新方法,并应用于解决长期争议的问题:Transformer中归一化层应置于何处?从小规模实验中恢复出大模型结果:预归一化在模型变大时表现更优。有了正确工具和理解,小规模实验终能兑现缩放定律的承诺。
原文摘要 · Abstract (English)
Scaling laws promised cost-effective experiments; six years later, they have yet to fully deliver. Instead, researchers have found them unreliable at small scales (starting at 4M parameters) and concluded that sizable models cannot be avoided. We show this is not the case: the confounding factor is hyperparameters. Small models are highly sensitive, but hyperparameter sensitivity fades with scale. This small-scale sensitivity makes scaling laws easy to miss because they only emerge on the fully tuned frontier, and reaching that frontier requires an extensive search far beyond what most ever run. By ablating the basic scaling law recipe, we show well-tuned hyperparameters matter more than any other ingredient. Further, we reveal why those hyperparameters become easier to find: as scale increases, the hyperparameter loss surface becomes lower dimensional. Nevertheless while scaling laws exist in small models, extrapolation hits statistical limitations. A holistic approach is required. Synthesizing our insights with the recent literature, we develop a new methodology for model-centric research and demonstrate it on a question that once took the field years to settle: where to place normalization layers in the transformer architecture. From small-scale experiments, we recover the large scale result: pre-normalization works better as models grow in size. With the right tools and a better understanding, small-scale experiments can deliver on scaling laws' long-awaited promise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。