arXiv:2605.20196cs.CLcs.AI2026-05

数据规模增长实质是逐步覆盖预测贡献谱的前沿。

Data Scaling as Progressive Coverage of a Predictive Contribution Spectrum

论文配图:Data Scaling as Progressive Coverage of a Predictive Contribution Spectrum
图 1 · 摘自论文原文
  • 用后缀自动机构建文本的全局KL预测贡献谱,量化每状态贡献。
  • 训练规模与预测谱尾部质量强相关,有效截断秩对数随数据量近似线性增长。
  • 适合研究模型缩放规律、数据效率或理论机器学习的研究者。

我们提出假设:真实数据的缩放规律由潜在预测贡献谱的渐进覆盖驱动,而非仅由词元频率尾部决定。基于后缀自动机表示文本语料,定义数据内在的全局KL预测贡献谱,其中每个状态贡献为其经验概率与相对于全局下一个词基线的KL偏差乘积。在12个真实语料上,该谱的尾部斜率已与固定小型GPT学习器的实证数据缩放指数高度相关。进一步地,针对每个训练规模N,通过匹配观测到的过损失与预处理的1000k全局KL谱的剩余尾部质量,定义有效截断秩K(N)。实验显示,log K与log N接近线性关系,原始谱的联合R²约0.96,平滑谱约0.90。这些结果为简单机制提供了强实证支持:训练规模推动有效前沿穿越预测状态谱,而该谱的剩余尾部质量可追踪剩余过损失。

原文摘要 · Abstract (English)

We investigate the hypothesis that real-data scaling laws are governed by progressive coverage of a latent predictive contribution spectrum rather than by token-frequency tails alone. We work with a suffix-automaton representation of text corpora and define a data-intrinsic global-KL predictive contribution spectrum, in which each state contributes according to its empirical mass times its KL deviation from a global next-token baseline. Across 12 real corpora, the tail slope of this spectrum is already strongly correlated with the empirical data-scaling exponent of a fixed small GPT learner. We then go beyond slope correlation and define, for each training size N, an effective truncation rank K(N) by matching the observed excess loss to the residual tail mass of the prepared 1000k global-KL spectrum. Empirically, log K is close to linear in log N, with pooled R^2 about 0.96 for the raw spectrum and R^2 about 0.90 for the smoothed spectrum. These findings provide strong empirical support for a simple mechanism picture: training scale advances an effective frontier through a predictive state spectrum, and the residual tail mass of that spectrum tracks the remaining excess loss.

数据缩放预测贡献模型性能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。