arXiv:2504.11393cs.LGcs.CL2025-04ICML被引 39

用小规模实验预测大模型最佳训练数据,准确率超80%。

DataDecide: How to Predict Best Pretraining Data with Small Experiments

  • 基于150M参数的小模型排名可预测1B参数大模型表现
  • 仅用0.01%算力,连续似然指标在多个基准上预测准确率>80%
  • 适合想低成本筛选预训练数据的研究者和工程师

由于大语言模型在不同数据集上预训练成本高昂,通过小规模实验来决策数据至关重要。我们构建了DataDecide——目前最全面的跨数据与规模的开源模型套件,涵盖25个语料库,覆盖不同来源、去重和过滤策略,最大达1000亿词元,模型规模最高10亿参数,3个随机种子。实验发现,单一小型规模(如1.5亿参数)下的模型排名是预测10亿参数目标规模下最优模型的强基线,正确率达约80%。在8种基准方法中,无一超越单尺度预测的算力-决策边界;但DataDecide可衡量未来缩放定律的改进。此外,使用连续似然指标作为小规模代理,在包含MMLU、ARC、HellaSwag、MBPP和HumanEval的基准上,仅需0.01%算力即可实现超过80%的预测准确率。

原文摘要 · Abstract (English)

Because large language models are expensive to pretrain on different datasets, using smaller-scale experiments to decide on data is crucial for reducing costs. Which benchmarks and methods of making decisions from observed performance at small scale most accurately predict the datasets that yield the best large models? To empower open exploration of this question, we release models, data, and evaluations in DataDecide -- the most extensive open suite of models over differences in data and scale. We conduct controlled pretraining experiments across 25 corpora with differing sources, deduplication, and filtering up to 100B tokens, model sizes up to 1B parameters, and 3 random seeds. We find that the ranking of models at a single, small size (e.g., 150M parameters) is a strong baseline for predicting best models at our larger target scale (1B) (~80% of com parisons correct). No scaling law methods among 8 baselines exceed the compute-decision frontier of single-scale predictions, but DataDecide can measure improvement in future scaling laws. We also identify that using continuous likelihood metrics as proxies in small experiments makes benchmarks including MMLU, ARC, HellaSwag, MBPP, and HumanEval >80% predictable at the target 1B scale with just 0.01% of the compute.

预训练数据小规模实验模型评估算力优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。