arXiv:2001.08361cs.LGstat.ML2020-01被引 8.9k

大模型更省数据,用少量数据训大模型最省算力。

Scaling Laws for Neural Language Models

论文配图:Scaling Laws for Neural Language Models
图 1 · 摘自论文原文
  • 用幂律关系建模模型大小、数据量和算力对损失的影响。
  • 大模型在相同算力下比小模型少用90%以上数据。
  • 适合追求算力效率的模型训练者和资源有限的研究者。

我们研究了语言模型在交叉熵损失上的经验缩放规律。损失随模型规模、数据集规模和训练算力呈幂律变化,部分趋势跨越七阶以上数量级。网络宽度或深度等其他架构细节在宽范围内影响甚微。简单的公式描述了过拟合与模型/数据规模的关系,以及训练速度与模型规模的关系。这些关系使我们能确定固定算力预算下的最优分配方式。大模型显著更高效:最优算力效率的训练策略是用相对较少的数据训练超大规模模型,并在远未收敛时停止。

原文摘要 · Abstract (English)

We study empirical scaling laws for language model performance on the cross-entropy loss. The loss scales as a power-law with model size, dataset size, and the amount of compute used for training, with some trends spanning more than seven orders of magnitude. Other architectural details such as network width or depth have minimal effects within a wide range. Simple equations govern the dependence of overfitting on model/dataset size and the dependence of training speed on model size. These relationships allow us to determine the optimal allocation of a fixed compute budget. Larger models are significantly more sample-efficient, such that optimally compute-efficient training involves training very large models on a relatively modest amount of data and stopping significantly before convergence.

语言模型缩放定律算力效率样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。