Ling 2.0用稀疏激活实现万亿参数语言模型高效推理
Every Activation Boosted: Scaling General Reasoner to 1 Trillion Open Language Foundation
- 采用高稀疏MoE架构,通过MTP提升推理效率
- 1万亿参数模型比密集模型节省7倍计算量
- 适合追求高效推理的开源大模型研究者
我们提出Ling 2.0,一种以每个激活都增强推理能力为原则构建的推理导向语言基础模型。该系列在统一的混合专家(MoE)范式下,从数十亿扩展至一万亿参数,强调高稀疏性、跨尺度一致性与基于实证缩放律的效率。包含三个非思考型(指令)模型:Ling-mini-2.0、Ling-flash-2.0和Ling-1T,参数量分别为160亿至1万亿,相较密集模型最高实现7倍活跃计算效率提升。Ling 2.0在模型架构、预训练、后训练与基础设施上实现协同创新:采用高稀疏MoE结合多任务处理(MTP)实现高效推理,使用推理导向数据与中期思维链(CoT)激活,采用强化学习微调(DFT、Evo-CoT),并支持全规模FP8训练与细粒度异构流水线。在万亿规模下,Ling-1T在推理准确率与计算效率间建立新帕累托前沿,证明当稀疏激活与推理目标对齐时,可实现可扩展且高效的智能。整体上,Ling 2.0为未来推理与思维模型的发展提供了连贯、开放、高效的基石,包括基于同一基础构建的Ring系列。
原文摘要 · Abstract (English)
We introduce Ling 2.0, a series reasoning-oriented language foundation built upon the principle that every activation boosts reasoning capability. Designed to scale from tens of billions to one trillion parameters under a unified Mixture-of-Experts (MoE) paradigm, Ling 2.0 emphasizes high sparsity, cross-scale consistency, and efficiency guided by empirical scaling laws. The series includes three non-thinking (instruct) models - Ling-mini-2.0, Ling-flash-2.0, and Ling-1T - ranging from 16B to 1T total parameters and achieving up to 7-fold active-compute efficiency compared with dense counterparts. Ling 2.0 integrates coordinated innovations across model architecture, pre-training, post-training, and infrastructure: a high-sparsity MoE with MTP for efficient reasoning, reasoning-oriented data and mid-training CoT activation, reinforcement-based fine-tuning (DFT, Evo-CoT), and full-scale FP8 training with fine-grained heterogeneous pipelines. At the trillion scale, Ling-1T establishes a new Pareto frontier of reasoning accuracy versus computational efficiency, demonstrating that sparse activation, when properly aligned with reasoning objectives, enables scalable and efficient intelligence. Collectively, Ling 2.0 provides a coherent, open, and efficient foundation for advancing future reasoning and thinking models, including the Ring series built upon the same base.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。