arXiv:2602.07488cs.LGcs.AI2026-02被引 19

首次从语言统计特性推导出大模型的缩放定律,无需参数调优。

Deriving Neural Scaling Laws from the statistics of natural language

  • 基于语言中词元相关性与条件熵的衰减规律,建立理论模型。
  • 预测的缩放指数与GPT-2、LLaMA在两个数据集上的实测结果高度一致。
  • 适用于理解大模型在有限数据下的性能增长规律,适合研究者参考。

尽管实验性的神经网络缩放定律极大地推动了大规模机器学习的发展,但现有理论无法对任何现代大语言模型在任意自然语言数据集上训练时的缩放指数进行定量预测。本文首次在数据受限缩放定律情形下提供了这样的理论。我们识别出语言的两个关键统计特性:(i)词元对之间成对相关性随时间间隔的衰减;(ii)以当前上下文长度为条件时,下一个词元的条件熵的衰减。进一步,我们从第一性原理出发,推导出一个仅依赖于这些统计量的简单公式,可无自由参数、无需合成数据模型地预测数据受限下的神经网络缩放指数。该理论在从头训练GPT-2和LLaMA风格模型于两个性质不同的基准数据集TinyStories和WikiText上获得的实验测量缩放定律中表现出惊人吻合。

原文摘要 · Abstract (English)

Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no existing theory can quantitatively predict the exponents of these important laws for any modern LLM trained on any natural language dataset. We provide the first such theory in the case of data-limited scaling laws. We isolate two key statistical properties of language that alone can predict neural scaling exponents: (i) the decay of pairwise token correlations with time separation between token pairs, and (ii) the decay of the next-token conditional entropy with the length of the conditioning context. We further derive a simple formula in terms of these statistics that predicts data-limited neural scaling exponents from first principles without any free parameters or synthetic data models. Our theory exhibits a remarkable match with experimentally measured neural scaling laws obtained from training GPT-2 and LLaMA style models from scratch on two qualitatively different benchmarks, TinyStories and WikiText.

缩放定律语言模型统计建模大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。