arXiv:2409.06624cs.CLcs.AI2024-09被引 1

优化语言混合比例,让大模型中文能力显著提升。

A Practice of Post-Training on Llama-3 70B with Optimal Selection of Additional Language Mixture Ratio

  • 通过实验确定8B模型的最佳语言混合比与学习率组合。
  • 70B模型经微调后在中文、数学、编程等任务上性能提升。
  • 成果已部署至真实聊天系统,表现良好,适合工业应用。

大型语言模型(LLM)常需持续预训练(CPT)以获得陌生语言技能或适应新领域。然而CPT的高昂训练成本要求谨慎选择关键超参数,如额外语言或领域语料的混合比例(ALMR)。目前缺乏系统研究将最优混合比例与实际模型性能关联起来,也未解决实验缩放规律与全尺寸模型部署之间的差距。本文对Llama-3 8B和70B进行CPT,以增强其中文能力。通过系统探索8B模型上附加语言混合比(ALMR)与学习率(LR)的最优关联,确定了理想的实验设置。在充分调参并后续微调后,模型不仅在中文相关基准上表现更优,还在数学、编程及情感智能等特定领域实现提升。最终70B版本模型被部署于真实聊天系统,取得了满意效果。

原文摘要 · Abstract (English)

Large Language Models (LLM) often need to be Continual Pre-Trained (CPT) to obtain unfamiliar language skills or adapt to new domains. The huge training cost of CPT often asks for cautious choice of key hyper-parameters such as the mixture ratio of extra language or domain corpus. However, there is no systematic study that bridges the gap between the optimal mixture ratio and the actual model performance, and the gap between experimental scaling law and the actual deployment in the full model size. In this paper, we perform CPT on Llama-3 8B and 70B to enhance its Chinese ability. We study the optimal correlation between the Additional Language Mixture Ratio (ALMR) and the Learning Rate (LR) on the 8B size which directly indicates the optimal experimental setup. By thorough choice of hyper-parameter, and subsequent fine-tuning, the model capability is improved not only on the Chinese-related benchmark but also in some specific domains including math, coding, and emotional intelligence. We deploy the final 70B version of LLM on a real-life chat system which obtains satisfying performance.

大模型训练中文优化超参调优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。