剖析规模定律研究中的偏差,揭示关键细节如何影响结论可靠性。
(Mis)Fitting: A Survey of Scaling Laws
- 分析不同实验设置对规模定律拟合结果的影响。
- 50篇论文中45篇使用幂律模型,但多数缺乏可复现细节。
- 提出检查清单,提升规模定律研究的透明度与可重复性。
现代基础模型高度依赖规模定律来指导关键训练决策。研究人员常通过小规模训练实验,基于损失或任务性能与规模的关系外推最优架构和超参数配置。这一过程涉及多个可变因素:拟合方程、训练设置、优化方法等,均可能影响所得到的定律及其结论。我们讨论了先前研究在最优标记数与参数比等关键问题上得出的分歧结论,并通过自身分析揭示特定细节变化对规模研究结果的显著影响。此外,我们系统调研了超过50篇研究规模趋势的论文:其中45篇采用幂律量化趋势,但大多数未报告复现所需的关键信息。为缓解此问题,我们提出一份作者应参考的检查清单,以提升规模定律研究的严谨性与可重复性。
原文摘要 · Abstract (English)
Modern foundation models rely heavily on using scaling laws to guide crucial training decisions. Researchers often extrapolate the optimal architecture and hyper parameters settings from smaller training runs by describing the relationship between, loss, or task performance, and scale. All components of this process vary, from the specific equation being fit, to the training setup, to the optimization method. Each of these factors may affect the fitted law, and therefore, the conclusions of a given study. We discuss discrepancies in the conclusions that several prior works reach, on questions such as the optimal token to parameter ratio. We augment this discussion with our own analysis of the critical impact that changes in specific details may effect in a scaling study, and the resulting altered conclusions. Additionally, we survey over 50 papers that study scaling trends: while 45 of these papers quantify these trends using a power law, most under-report crucial details needed to reproduce their findings. To mitigate this, we we propose a checklist for authors to consider while contributing to scaling law research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。