用多样化合成方法打造1000亿词的高质量问答数据集,提升大模型知识能力
Large-Scale Diverse Synthesis for Mid-Training
- 从多源数据出发,用大模型分阶段生成跨学科高难度问题
- 中段训练使小模型在12个基准上平均提升12.74%,达到新顶尖水平
- 适合做模型知识增强、数据合成与评测的开发者和研究者
高质量、知识密集型训练数据的稀缺性制约了大语言模型的发展,传统语料库信息有限。已有研究虽尝试合成依赖语料的问答数据以提升模型性能,但在跨领域场景下仍面临可扩展性差与知识多样性不足的问题。为此,我们设计了学科与难度标注体系,揭示模型在理工科及高难度数据中的缺陷。提出全新多样化合成流程,构建包含1000亿词的大型问答数据集BoostQA。该框架:(1)从异构来源收集初始数据;(2)利用DeepSeek-R1实现聚焦理工科的多学段生成,增强多样性并缓解难度退化;(3)通过DeepSeek-V3对答案进行优化,提升输出质量。将BoostQA用于中段训练(介于预训练与后训练之间),显著提升领域知识获取能力与数据质量。实验表明,基于400亿词数据对Llama-3 8B进行中段训练,其在MMLU和CMMLU上平均提升12.74%,并在12个基准上达成当前最优表现。BoostQA具备强可扩展性,随模型规模、数据量与初始浮点运算量增长,性能持续提升。
原文摘要 · Abstract (English)
The scarcity of high-quality, knowledge-intensive training data hinders the development of large language models (LLMs), as traditional corpora provide limited information. Previous studies have synthesized and integrated corpora-dependent question-answering (QA) data to improve model performance but face challenges in QA data scalability and knowledge diversity, particularly in cross-domain contexts. Furthermore, leveraging our designed discipline and difficulty annotation system, we probe model deficiencies in STEM disciplines and high-difficulty data. To overcome these limitations, we propose a novel diversified pipeline to synthesize BoostQA, a 100B-token large-scale QA dataset. Our synthesis framework: (1) curates seed data from heterogeneous sources; (2) utilizes DeepSeek-R1 to implement STEM-focused multi-grade synthesis to boost data diversity and high-difficulty synthesis to mitigate difficulty degradation; (3) refines answers via DeepSeek-V3 to improve output quality. We utilize BoostQA in mid-training, a mid-stage between pre-training and post-training, to optimize domain-specific knowledge acquisition and enhance data quality. Our method enables Llama-3 8B, mid-trained on a 40B-token dataset, to achieve an average improvement of $\mathbf{12.74\%}$ on MMLU and CMMLU and establish SOTA average performance across 12 benchmarks. BoostQA also demonstrates robust scalability, with performance consistently improving as model size, data volume, and initial FLOPs scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。