梳理大模型中段训练方法,揭示其提升推理与编码能力的关键作用
A Survey on LLM Mid-Training
- 定义中段训练为连接预训练与后训练的独立阶段,聚焦能力增强
- 通过中间数据与策略优化,显著提升数学、编程、长文本处理能力
- 适合研究大模型训练流程与能力进化的学者参考
近年来基础模型的发展凸显了多阶段训练的优势,尤其强调中段训练作为连接预训练与后训练的重要环节。中段训练利用中间数据和计算资源,系统性地增强数学、编程、推理及长上下文处理等特定能力,同时保持基础性能。本文首次对大语言模型的中段训练进行正式定义,探讨涵盖数据筛选、训练策略与模型结构优化在内的优化框架。分析主流模型在目标驱动干预下的实现方式,说明中段训练在大模型能力渐进发展中的独特且关键作用。通过厘清其贡献,本综述构建了完整分类体系并提供可操作洞见,助力未来大模型研究与创新。
原文摘要 · Abstract (English)
Recent advances in foundation models have highlighted the significant benefits of multi-stage training, with a particular emphasis on the emergence of mid-training as a vital stage that bridges pre-training and post-training. Mid-training is distinguished by its use of intermediate data and computational resources, systematically enhancing specified capabilities such as mathematics, coding, reasoning, and long-context extension, while maintaining foundational competencies. This survey provides a formal definition of mid-training for large language models (LLMs) and investigates optimization frameworks that encompass data curation, training strategies, and model architecture optimization. We analyze mainstream model implementations in the context of objective-driven interventions, illustrating how mid-training serves as a distinct and critical stage in the progressive development of LLM capabilities. By clarifying the unique contributions of mid-training, this survey offers a comprehensive taxonomy and actionable insights, supporting future research and innovation in the advancement of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。