模型越大越强?新理论揭示训练时长与模型规模等效
Unified Neural Network Scaling Laws and Scale-time Equivalence
- 发现模型大小与训练时间可等效扩展,小模型久训可媲大模型
- 统一理论解释双下降、标签噪声敏感等反直觉现象
- 适合想高效训练大模型的研究者与工程师参考
随着神经网络持续增大而数据集增长有限,理解性能提升潜力变得至关重要:是扩大模型规模还是增加数据量更重要?神经网络缩放定律描述了测试误差随模型规模和数据量的变化关系,日益重要。然而现有缩放定律通常仅适用于有限范围,且难以解释双下降等经典现象。本文首次建立模型规模、训练时间和数据量三者相互作用的理论框架,提出模型规模与训练时间成比例扩展具有等效性。该尺度-时间等价性挑战了当前大模型短时训练的做法,表明小模型长期训练可达到同等效果,并由此提出从少量数据长期训练的小模型预测大规模模型性能的新方法。结合线性模型对双下降的分析,我们推导出统一的理论缩放定律,并在多个视觉基准和架构上验证。该定律解释了若干此前未解之谜:大模型对数据需求降低、过参数化模型对标签噪声更敏感,以及模型规模增大未必提升性能的现象。研究成果对实际部署神经网络具有重要意义,提供了更高效、低成本的训练路径。
原文摘要 · Abstract (English)
As neural networks continue to grow in size but datasets might not, it is vital to understand how much performance improvement can be expected: is it more important to scale network size or data volume? Thus, neural network scaling laws, which characterize how test error varies with network size and data volume, have become increasingly important. However, existing scaling laws are often applicable only in limited regimes and often do not incorporate or predict well-known phenomena such as double descent. Here, we present a novel theoretical characterization of how three factors -- model size, training time, and data volume -- interact to determine the performance of deep neural networks. We first establish a theoretical and empirical equivalence between scaling the size of a neural network and increasing its training time proportionally. Scale-time equivalence challenges the current practice, wherein large models are trained for small durations, and suggests that smaller models trained over extended periods could match their efficacy. It also leads to a novel method for predicting the performance of large-scale networks from small-scale networks trained for extended epochs, and vice versa. We next combine scale-time equivalence with a linear model analysis of double descent to obtain a unified theoretical scaling law, which we confirm with experiments across vision benchmarks and network architectures. These laws explain several previously unexplained phenomena: reduced data requirements for generalization in larger models, heightened sensitivity to label noise in overparameterized models, and instances where increasing model scale does not necessarily enhance performance. Our findings hold significant implications for the practical deployment of neural networks, offering a more accessible and efficient path to training and fine-tuning large models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。