arXiv:2508.10975cs.LGcs.CL2025-08被引 17

用合成数据突破万亿级预训练瓶颈,性能超越现有方法。

BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining

  • 构建新框架BeyondWeb,系统优化合成数据生成质量。
  • 30亿参数模型在合成数据上训练效果超80亿参数模型。
  • 适合追求高效高质量预训练的团队与研究者。

近期大语言模型预训练进展表明,单纯扩大数据量终将遭遇收益递减,陷入数据瓶颈。为此,利用合成数据进行预训练成为突破性能上限的可行路径。然而,影响合成数据质量的因素仍不明确。本文提出BeyondWeb,一种生成高质量合成数据的框架,显著扩展了传统网络规模数据集的能力。在14项基准测试中,其平均表现优于当前最优合成数据集Cosmopedia和Nemotron-CC高质量子集(Nemotron-Synth)达5.1个百分点和2.6个百分点。训练速度较开源网络数据快7.7倍,较Nemotron-Synth快2.7倍。值得注意的是,一个30亿参数模型在1800亿词元的BeyondWeb数据上训练,性能超过在Cosmopedia上训练的80亿参数模型。此外,我们总结了关于合成数据生成的关键洞见:驱动优势的因素、应重写的数据类型及方式,以及模型规模与架构对数据质量的影响。整体表明,高质量合成数据无万能解法,需协同优化多重因素,仅靠简单方法难以获得显著提升,而精心设计的方法可带来质变,如BeyondWeb所示。

原文摘要 · Abstract (English)

Recent advances in large language model (LLM) pretraining have shown that simply scaling data quantity eventually leads to diminishing returns, hitting a data wall. In response, the use of synthetic data for pretraining has emerged as a promising paradigm for pushing the frontier of performance. Despite this, the factors affecting synthetic data quality remain poorly understood. In this work, we introduce BeyondWeb, a synthetic data generation framework that produces high-quality synthetic data for pretraining. BeyondWeb significantly extends the capabilities of traditional web-scale datasets, outperforming state-of-the-art synthetic pretraining datasets such as Cosmopedia and Nemotron-CC's high-quality synthetic subset (Nemotron-Synth) by up to 5.1 percentage points (pp) and 2.6pp, respectively, when averaged across a suite of 14 benchmark evaluations. It delivers up to 7.7x faster training than open web data and 2.7x faster than Nemotron-Synth. Remarkably, a 3B model trained for 180B tokens on BeyondWeb outperforms an 8B model trained for the same token budget on Cosmopedia. We also present several insights from BeyondWeb on synthetic data for pretraining: what drives its benefits, which data to rephrase and how, and the impact of model size and family on data quality. Overall, our work shows that there's no silver bullet for generating high-quality synthetic pretraining data. The best outcomes require jointly optimizing many factors, a challenging task that requires rigorous science and practical expertise. Naive approaches can yield modest improvements, potentially at great cost, while well-executed methods can yield transformative improvements, as exemplified by BeyondWeb.

合成数据预训练大模型效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。