通过系统实验发现,结构化重述能显著提升合成数据质量。
How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
- 用表格、问答等结构化格式重写网页文本,效果优于传统方法。
- 生成模型超过10亿参数后性能不再提升,小模型即可满足需求。
- 精选原始数据源比复杂提示设计更重要,适合大模型预训练数据构建者。
合成数据是训练大语言模型的标准组件,但关于重述策略、生成模型和源数据等设计维度的系统性比较仍缺失。我们进行了大规模受控实验,生成超一万亿标记符,识别出将网络文本转化为合成预训练数据的关键因素。结果表明,表格、数学题、问答对和教程等结构化输出格式始终优于人工筛选的网络基线及以往合成方法。值得注意的是,生成模型规模超过10亿参数后不再带来性能增益。分析还显示,原始数据混合选择对性能影响显著。基于此,我们构建了 extbf{ extsc{FinePhrase}},一个包含4860亿标记符的开放重述网页文本数据集。实验证明, extsc{FinePhrase} 在所有现有合成数据基线上表现更优,同时生成成本降低达30倍。我们已向研究社区公开该数据集、全部提示词与生成框架。
原文摘要 · Abstract (English)
Synthetic data is a standard component in training large language models, yet systematic comparisons across design dimensions, including rephrasing strategy, generator model, and source data, remain absent. We conduct extensive controlled experiments, generating over one trillion tokens, to identify critical factors in rephrasing web text into synthetic pretraining data. Our results reveal that structured output formats, such as tables, math problems, FAQs, and tutorials, consistently outperform both curated web baselines and prior synthetic methods. Notably, increasing the size of the generator model beyond 1B parameters provides no additional benefit. Our analysis also demonstrates that the selection of the original data used for mixing substantially influences performance. By applying our findings, we develop \textbf{\textsc{FinePhrase}}, a 486-billion-token open dataset of rephrased web text. We show that \textsc{FinePhrase} outperforms all existing synthetic data baselines while reducing generation costs by up to 30 times. We provide the dataset, all prompts, and the generation framework to the research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。