arXiv:2504.07866cs.CLcs.AI2025-04被引 17

1350亿参数大模型在昇腾芯片上实现高效训练,性能超越多个主流模型。

Pangu Ultra: Pushing the Limits of Dense Large Language Models on Ascend NPUs

  • 提出深度缩放夹心归一化,解决深层模型训练中的损失突增问题。
  • 使用8192个昇腾NPU训练,基于13.2万亿高质量文本,推理能力显著提升。
  • 证明昇腾平台可高效训练超百亿参数稠密模型,适合工业级应用。

我们提出 Pangu Ultra,一个拥有1350亿参数、基于密集Transformer模块并在昇腾神经处理单元(Ascend NPUs)上训练的大语言模型。尽管近年来大语言模型在规模与能力上取得空前进展,但训练如此大规模模型仍面临显著的优化与系统挑战。为稳定训练过程,我们提出深度缩放夹心归一化(depth-scaled sandwich normalization),有效消除深层模型训练中的损失峰值。我们在13.2万亿多样且高质量的文本数据上预训练该模型,并在后训练阶段进一步增强其推理能力。为实现高效的大规模训练,我们利用8192个昇腾NPU并进行一系列系统优化。在多个多样化基准上的评估表明,Pangu Ultra显著提升了稠密大语言模型的性能,超越Llama 405B和Mistral Large 2,甚至达到与参数更多但结构稀疏的DeepSeek-R1相竞争的水平。我们的探索证明,昇腾NPU能够高效、有效地训练超过1000亿参数的稠密模型。模型与系统将面向商业客户开放。

原文摘要 · Abstract (English)

We present Pangu Ultra, a Large Language Model (LLM) with 135 billion parameters and dense Transformer modules trained on Ascend Neural Processing Units (NPUs). Although the field of LLM has been witnessing unprecedented advances in pushing the scale and capability of LLM in recent years, training such a large-scale model still involves significant optimization and system challenges. To stabilize the training process, we propose depth-scaled sandwich normalization, which effectively eliminates loss spikes during the training process of deep models. We pre-train our model on 13.2 trillion diverse and high-quality tokens and further enhance its reasoning capabilities during post-training. To perform such large-scale training efficiently, we utilize 8,192 Ascend NPUs with a series of system optimizations. Evaluations on multiple diverse benchmarks indicate that Pangu Ultra significantly advances the state-of-the-art capabilities of dense LLMs such as Llama 405B and Mistral Large 2, and even achieves competitive results with DeepSeek-R1, whose sparse model structure contains much more parameters. Our exploration demonstrates that Ascend NPUs are capable of efficiently and effectively training dense models with more than 100 billion parameters. Our model and system will be available for our commercial customers.

大模型昇腾稠密模型训练优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。