大模型训练中段优化新范式,提升性能与稳定性。
Mid-Training of Large Language Models: A Survey
- 引入中段训练阶段,分阶段优化数据、学习率和上下文长度。
- 实测可缓解噪声词导致的性能下降,提升模型泛化能力。
- 适合大模型研发者参考,推动训练流程标准化。
大语言模型通常通过大规模预训练后进行任务微调来构建。近期进展表明,引入中间的中段训练阶段至关重要:该阶段通过多个渐进式调整阶段,提升数据质量、优化学习率调度并扩展上下文长度。这一过程能有效缓解噪声词带来的收益递减问题,稳定收敛过程,并增强模型在训练后期的能力。其有效性可通过梯度噪声尺度、信息瓶颈理论与课程学习机制共同解释,有助于提升模型泛化性与抽象能力。尽管已在先进系统中广泛采用,但此前尚无将中段训练作为统一范式的综述。本文首次构建了涵盖数据分布、学习率调度与长上下文扩展三方面的分类体系,提炼实用经验,建立评估基准并报告性能增益,支持跨模型的结构化比较。同时指出开放挑战,提出未来研究与实践方向。
原文摘要 · Abstract (English)
Large language models (LLMs) are typically developed through large-scale pre-training followed by task-specific fine-tuning. Recent advances highlight the importance of an intermediate mid-training stage, where models undergo multiple annealing-style phases that refine data quality, adapt optimization schedules, and extend context length. This stage mitigates diminishing returns from noisy tokens, stabilizes convergence, and expands model capability in late training. Its effectiveness can be explained through gradient noise scale, the information bottleneck, and curriculum learning, which together promote generalization and abstraction. Despite widespread use in state-of-the-art systems, there has been no prior survey of mid-training as a unified paradigm. We introduce the first taxonomy of LLM mid-training spanning data distribution, learning-rate scheduling, and long-context extension. We distill practical insights, compile evaluation benchmarks, and report gains to enable structured comparisons across models. We also identify open challenges and propose avenues for future research and practice.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。