TeleChat3-MoE实现千亿至万亿参数模型高效训练,突破大规模语言模型部署瓶颈。
Training Report of TeleChat3-MoE
- 采用专家混合(MoE)架构,支持千亿到万亿参数模型端到端训练。
- 在数千设备集群上实现近线性扩展,吞吐量显著提升。
- 适用于需要超大规模语言模型训练的科研与工业场景。
TeleChat3-MoE 是最新一代基于专家混合(Mixture-of-Experts, MoE)架构的 TeleChat 大语言模型系列,参数量达 1050 亿至超过一万亿,依托 Ascend NPU 集群实现端到端训练。本技术报告重点阐述支撑前沿模型规模可靠高效扩展的底层训练基础设施。我们提出逐算子与端到端数值精度验证方法,确保跨硬件平台与分布式并行策略的一致性。引入多项性能优化:交错流水线调度、面向注意力机制的长序列数据调度、分层重叠通信以支持专家并行,以及基于 DVM 的算子融合。此外,提出基于解析估计与整数线性规划的系统性并行化框架,用于优化多维并行配置。还针对大规模训练中的主机与设备瓶颈,提出集群级优化方法。这些改进使模型在包含数千设备的集群上实现显著吞吐提升与近线性扩展,为硬件生态上的大模型研发提供了坚实基础。
原文摘要 · Abstract (English)
TeleChat3-MoE is the latest series of TeleChat large language models, featuring a Mixture-of-Experts (MoE) architecture with parameter counts ranging from 105 billion to over one trillion,trained end-to-end on Ascend NPU cluster. This technical report mainly presents the underlying training infrastructure that enables reliable and efficient scaling to frontier model sizes. We detail systematic methodologies for operator-level and end-to-end numerical accuracy verification, ensuring consistency across hardware platforms and distributed parallelism strategies. Furthermore, we introduce a suite of performance optimizations, including interleaved pipeline scheduling, attention-aware data scheduling for long-sequence training,hierarchical and overlapped communication for expert parallelism, and DVM-based operator fusion. A systematic parallelization framework, leveraging analytical estimation and integer linear programming, is also proposed to optimize multi-dimensional parallelism configurations. Additionally, we present methodological approaches to cluster-level optimizations, addressing host- and device-bound bottlenecks during large-scale training tasks. These infrastructure advancements yield significant throughput improvements and near-linear scaling on clusters comprising thousands of devices, providing a robust foundation for large-scale language model development on hardware ecosystems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。