大模型时代,旧有正则化原则失效,需新范式指导模型扩展。
Rethinking Conventional Wisdom in Machine Learning: From Generalization to Scaling
- 用大规模训练替代正则化,突破传统泛化思维
- 发现缩放定律交叉现象,小规模有效方法未必适用于大模型
- 提出两个核心问题:如何指导模型缩放?如何在极限规模下比较模型?
大规模语言模型预训练的成功与缩放定律的发现标志着机器学习范式的转变。核心目标已从最小化泛化误差转向降低近似误差,最优策略也从广义正则化转向模型扩展。这引发关键问题:过去在泛化主导时代有效的原则,在当前以缩放为中心的大模型时代是否依然成立?本文检验了几项基于正则化的经典原则在大模型时代的适用性,包括显式的L2正则化、小批次和大学习率带来的隐式正则化。此外,我们发现了新的“缩放定律交叉”现象——两条缩放曲线在特定规模下相交,表明小规模有效的技术可能无法推广至更大规模。这些发现凸显了新范式下的两大根本问题:第一,若正则化不再是主要指导原则,新的缩放引导原则是什么?第二,当仅能进行一次实验时,如何可靠地比较大规模模型?
原文摘要 · Abstract (English)
The remarkable success of large language pretraining and the discovery of scaling laws signify a paradigm shift in machine learning. Notably, the primary objective has evolved from minimizing generalization error to reducing approximation error, and the most effective strategy has transitioned from regularization (in a broad sense) to scaling up models. This raises a critical question: Do the established principles that proved successful in the generalization-centric era remain valid in this new era of scaling? This paper examines several influential regularization-based principles that may no longer hold true in the scaling-centric, large language model (LLM) era. These principles include explicit L2 regularization and implicit regularization through small batch sizes and large learning rates. Additionally, we identify a new phenomenon termed ``scaling law crossover,'' where two scaling curves intersect at a certain scale, implying that methods effective at smaller scales may not generalize to larger ones. Together, these observations highlight two fundamental questions within this new paradigm: $\bullet$ Guiding Principles for Scaling: If regularization is no longer the primary guiding principle for model design, what new principles are emerging to guide scaling? $\bullet$ Model Comparison at Scale: How to reliably and effectively compare models at the scale where only a single experiment is feasible?
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。