早加推理数据能显著提升大模型能力,比后期训练更有效。
Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data
- 将推理数据提前至预训练阶段,效果优于后期微调。
- 预训练阶段引入多样推理数据可提升11%,微调阶段则需高质量数据(+15%)。
- 过量微调数据会抵消早期推理数据的好处,需合理分配资源。
当前主流方法是通过高质量推理数据进行后训练来增强大模型的推理能力。然而,已有研究表明推理数据也逐渐被引入中期训练阶段——这一做法较为私密且公开描述较少。多数前沿模型的预训练语料不透明,导致推理数据在不同训练阶段的影响尚不明确。本研究首次系统性地探讨了在不同训练阶段引入推理数据(规模、多样性、质量各异)对模型性能的影响。结果表明:将推理数据前置到预训练阶段至关重要,平均可提升19%,其建立的基础能力无法仅靠后期SFT完全复制,即使使用更多数据亦然。我们发现最优数据分配具有非对称性:预训练阶段最受益于推理模式的广泛多样性(+11%),而微调阶段则更依赖数据质量(+15%)。高质预训练数据具有潜在效应,仅在微调后才被激活;盲目扩大微调数据反而有害,会冲淡早期推理注入的优势。研究挑战了语言建模与推理分离的传统观念,为全训练流程的数据战略分配提供原则性指导。
原文摘要 · Abstract (English)
The prevailing paradigm for enhancing the reasoning abilities of LLMs revolves around post-training on high-quality, reasoning-intensive data. While emerging literature suggests that reasoning data is increasingly incorporated also during the mid-training stage-a practice that is relatively more proprietary and less openly characterized-the role of such data in pretraining remains unclear. In particular, due to the opaqueness of pretraining corpora in most frontier models, the effect of reasoning data introduced at different phases of pre- and/or post-training is relatively less reported in the scientific literature. This raises several important questions: Is adding reasoning data earlier during pretraining any better than introducing it during post-training? Could earlier inclusion risk overfitting and harm generalization, or instead establish durable foundations that later fine-tuning cannot recover? We conduct the first systematic study of how reasoning data-varying in scale, diversity, and quality-affects LLM performance when introduced at different stages of training. We find that front-loading reasoning data into pretraining is critical (19% avg gain), establishing foundational capabilities that cannot be fully replicated by later-stage SFT, even with more data. We uncover an asymmetric principle for optimal data allocation: pretraining benefits most from broad diversity in reasoning patterns (11% avg gain), while SFT is more sensitive to data quality (15% avg gain). We show that high-quality pretraining data has latent effects, activated only after SFT, and that naively scaling SFT data can be detrimental, washing away the benefits of early reasoning injection. Our results challenge the conventional separation of language modeling and reasoning, providing a principled guide for strategically allocating data across the entire training pipeline to build more capable models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。