arXiv:2510.23081cs.CL2025-10综述被引 19

梳理大模型中段训练方法,揭示其提升推理与编码能力的关键作用

A Survey on LLM Mid-Training

  • 定义中段训练为连接预训练与后训练的独立阶段,聚焦能力增强
  • 通过中间数据与策略优化,显著提升数学、编程、长文本处理能力
  • 适合研究大模型训练流程与能力进化的学者参考

近年来基础模型的发展凸显了多阶段训练的优势,尤其强调中段训练作为连接预训练与后训练的重要环节。中段训练利用中间数据和计算资源,系统性地增强数学、编程、推理及长上下文处理等特定能力,同时保持基础性能。本文首次对大语言模型的中段训练进行正式定义,探讨涵盖数据筛选、训练策略与模型结构优化在内的优化框架。分析主流模型在目标驱动干预下的实现方式,说明中段训练在大模型能力渐进发展中的独特且关键作用。通过厘清其贡献,本综述构建了完整分类体系并提供可操作洞见,助力未来大模型研究与创新。

原文摘要 · Abstract (English)

Recent advances in foundation models have highlighted the significant benefits of multi-stage training, with a particular emphasis on the emergence of mid-training as a vital stage that bridges pre-training and post-training. Mid-training is distinguished by its use of intermediate data and computational resources, systematically enhancing specified capabilities such as mathematics, coding, reasoning, and long-context extension, while maintaining foundational competencies. This survey provides a formal definition of mid-training for large language models (LLMs) and investigates optimization frameworks that encompass data curation, training strategies, and model architecture optimization. We analyze mainstream model implementations in the context of objective-driven interventions, illustrating how mid-training serves as a distinct and critical stage in the progressive development of LLM capabilities. By clarifying the unique contributions of mid-training, this survey offers a comprehensive taxonomy and actionable insights, supporting future research and innovation in the advancement of LLMs.

大模型中段训练能力增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。