arXiv:2607.25271cs.LGcs.AI2026-07

提出新框架,解决模型训练中算力与数据不匹配问题。

Bridging Compute- and Data-Optimal Pretraining

论文配图:Bridging Compute- and Data-Optimal Pretraining
图 1 · 摘自论文原文
  • 引入令牌有效性函数η,衡量重复或改写数据的价值
  • 发现数据扩展效果随模型大小和数据量饱和,存在边际递减
  • 揭示三类训练状态,指出传统算力最优策略多不适用

经典算力最优缩放定律假设可无限获取新鲜预训练数据,但当前预训练正进入算力增速超过高质量数据供给的阶段。本文提出计算-数据(CD)缩放定律,统一算力最优(数据随算力自由扩展)与数据最优(数据集固定而算力无限制)两种情形。该框架引入令牌有效性函数η,量化通过多轮重复或改写生成的衍生令牌相对于原始令牌的价值,取值范围从完全替代到无价值。我们在14M至600M参数模型上,基于Dolma-3语料库对多轮重复和改写两种数据扩展策略拟合η。结果表明,令牌有效性并非恒定,而是同时依赖于模型规模、每参数令牌数及衍生数据量,并在语料库扩展时趋于饱和。η的函数形式显示,随着模型规模或数据可用性增加,以算力替代数据的收益递减。该框架还划分出三类训练阶段——算力受限、数据受限、模型受限,并表明经典算力最优分配在大多数实际场景下并非最优。

原文摘要 · Abstract (English)

Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a regime in which compute grows faster than the availability of high-quality data. We propose Compute-Data (CD) scaling laws, a unified framework that bridges compute-optimal scaling, where data scales freely with compute, and data-optimal scaling, where the corpus is fixed while compute can grow without bound. CD scaling extends classical scaling laws by introducing a token-effectiveness function, $η$, which quantifies the value of a derived token-produced, for example, through multi-epoch repetition or paraphrasing-relative to a fresh token, ranging from a perfect substitute to having no value. We fit $η$ for two data-expansion strategies, multi-epoch repetition and paraphrasing, across model sizes from 14M to 600M parameters using the Dolma-3 corpus. We find that token effectiveness is far from constant: it depends jointly on model size, the tokens-per-parameter ratio, and the amount of derived data, and it saturates as the corpus is expanded. The functional form of $η$ implies diminishing returns when substituting compute for data as either model size or data availability increases. It also partitions training into three operational regimes---compute-bound, data-bound, and model-bound---and shows that classical compute-optimal allocation is suboptimal across most practically relevant settings.

缩放定律数据效率预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。