arXiv:2505.12082cs.CLcs.LG2025-05NeurIPS被引 32

通过合并预训练检查点,显著提升大模型性能并降低训练成本。

Model Merging in Pre-training of Large Language Models

  • 合并恒定学习率训练的模型检查点,提升性能。
  • 可准确预测退火行为,实现更高效训练。
  • 适合追求低成本高效训练的开发者参考。

模型合并已成为提升大语言模型的有前景技术,但在大规模预训练中的应用仍相对未被探索。本文全面研究了预训练过程中模型合并技术。通过对从数百万到超过1000亿参数的密集模型与专家混合(MoE)架构进行大量实验,我们发现:以恒定学习率训练的检查点进行合并,不仅能显著提升性能,还能准确预测退火行为。这些改进使模型开发更高效,训练成本大幅降低。我们对合并策略和超参数的详细消融研究,揭示了其内在机制,并发现了新应用场景。通过全面实验分析,我们为开源社区提供了有效的模型合并预训练实践指南。

原文摘要 · Abstract (English)

Model merging has emerged as a promising technique for enhancing large language models, though its application in large-scale pre-training remains relatively unexplored. In this paper, we present a comprehensive investigation of model merging techniques during the pre-training process. Through extensive experiments with both dense and Mixture-of-Experts (MoE) architectures ranging from millions to over 100 billion parameters, we demonstrate that merging checkpoints trained with constant learning rates not only achieves significant performance improvements but also enables accurate prediction of annealing behavior. These improvements lead to both more efficient model development and significantly lower training costs. Our detailed ablation studies on merging strategies and hyperparameters provide new insights into the underlying mechanisms while uncovering novel applications. Through comprehensive experimental analysis, we offer the open-source community practical pre-training guidelines for effective model merging.

模型合并预训练大模型效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。