arXiv:2605.02364cs.CL2026-05被引 1

提出信息尺度定律,精准预测大模型训练性能

InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition

论文配图:InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition
图 1 · 摘自论文原文
  • 将预训练视为信息积累过程,融合数据质量与重复度影响
  • 在70亿至4250亿令牌规模上误差低于1%,预测精度高
  • 适用于算力受限时的数据配方优化,适合模型训练调优者

在数据量有限的场景下,提高高质量数据权重虽能提升性能,但过度加权会加剧重复,导致性能下降。标准尺度定律无法可靠外推不同数据混合方案或重复情况下的表现,使大规模训练下的数据配方选择缺乏依据。为此,我们提出InfoLaw(信息尺度定律),一个基于数据的尺度建模框架,可预测损失值,输入包括消耗的令牌数、模型规模、数据混合权重和重复率。核心思想是将预训练建模为信息累积过程:数据质量决定信息密度,重复则引发规模依赖的边际收益递减。我们收集了在不同规模、质量分布和重复水平数据集上训练后的模型性能,构建信息模型以准确预测实际表现。InfoLaw在未见过的数据配方和更大规模训练(最高达70亿参数、4250亿令牌)中表现出色,平均绝对误差仅0.15%,最大误差0.96%,且能稳定外推至过拟合场景,实现不同算力预算下的高效数据配方选择。

原文摘要 · Abstract (English)

Upweighting high-quality data in LLM pretraining often improves performance, but in datalimited regimes, especially under overtraining, stronger upweighting increases repetition and can degrade performance. However, standard scaling laws do not reliably extrapolate across mixture recipes or under repetitions, making the selection for optimal data recipes at scaling underdetermined. To solve this, we introduce InfoLaw (Information Scaling Laws), a data-aware scaling framework that predicts loss from consumed tokens, model size, data mixture weights, and repetition. The key idea is to model pretraining as information accumulation, where quality controls information density and repetition induces scaledependent diminishing returns. We first collect the model performance after training on datasets that vary in scale, quality distribution, and repetition level. Then we build up the modeling for information so that information accurately predicts those model performance. InfoLaw predicts performance on unseen data recipes and larger scale runs (up to 7B, 425B tokens) with 0.15% mean and 0.96% max absolute error in loss, and it extrapolates reliably across overtraining levels, enabling efficient data-recipe selection under varying compute budgets.

大模型训练数据质量尺度定律信息建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。