arXiv:2601.10684cs.LGcond-mat.dis-nn2026-01被引 12

即使数据无幂律结构,模型仍出现缩放规律,揭示其根源可能在模型本身而非数据。

On the origin of neural scaling laws: from random graphs to natural language

  • 用随机图上的随机游走训练Transformer,验证无幂律数据也能产生缩放规律
  • 语言复杂度逐级降低时,缩放指数单调变化,揭示规律与任务复杂度相关
  • 2层模型即可复现经典语言建模缩放结果,为高效训练提供新思路

缩放定律在现代AI发展中起关键作用,可预测模型性能随数据量、算力和参数量增加的变化。学界普遍认为其源于数据中固有的幂律结构。本文研究在可调复杂度的图上训练Transformer预测随机游走(双词)时的缩放行为,发现即使数据相关性中无幂律结构,仍能出现神经网络缩放定律。进一步通过从4层、2层、1层Transformer到语言双词模型逐步简化自然语言生成过程,观察到缩放指数的单调演化。实验还涵盖基于Erdös-Renyi与尺度不变的Barabási-Albert随机图的随机游走。最后重新审视常规语言建模缩放规律,表明使用2层变压器(上下文长度50)即可重现多个核心结果;批判了以往文献中的多种拟合方法,提出一种新的计算最优曲线获取方式,并初步显示最大更新参数化可能比标准参数化更高效。

原文摘要 · Abstract (English)

Scaling laws have played a major role in the modern AI revolution, providing practitioners predictive power over how the model performance will improve with increasing data, compute, and number of model parameters. This has spurred an intense interest in the origin of neural scaling laws, with a common suggestion being that they arise from power law structure already present in the data. In this paper we study scaling laws for transformers trained to predict random walks (bigrams) on graphs with tunable complexity. We demonstrate that this simplified setting already gives rise to neural scaling laws even in the absence of power law structure in the data correlations. We further consider dialing down the complexity of natural language systematically, by training on sequences sampled from increasingly simplified generative language models, from 4,2,1-layer transformer language models down to language bigrams, revealing a monotonic evolution of the scaling exponents. Our results also include scaling laws obtained from training on random walks on random graphs drawn from Erdös-Renyi and scale-free Barabási-Albert ensembles. Finally, we revisit conventional scaling laws for language modeling, demonstrating that several essential results can be reproduced using 2 layer transformers with context length of 50, provide a critical analysis of various fits used in prior literature, demonstrate an alternative method for obtaining compute optimal curves as compared with current practice in published literature, and provide preliminary evidence that maximal update parameterization may be more parameter efficient than standard parameterization.

缩放定律Transformer语言建模随机图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。