arXiv:2606.15079cs.CLcs.AI2026-06被引 3

超大规模智能体模型实现低延迟与强推理兼得,开源全系列参数

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale

论文配图:Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale
图 1 · 摘自论文原文
  • 通过架构迁移与大规模微调,升级基础模型提升效率与能力
  • 混合线性注意力设计支持长上下文训练与解码,推理速度更快
  • 提出KPop强化学习框架,支持复杂智能体任务的稳定训练

高效可扩展的智能体智能需要在低延迟响应与强推理能力之间取得平衡,同时保持训练、服务与部署的实用性。本文提出Ling-2.6与Ring-2.6系列模型,针对这一挑战在万亿参数规模上实现突破。Ling-2.6优化即时响应与每输出词元的高能力表现,Ring-2.6则专为深度推理与高级智能体工作流设计。模型通过架构迁移预训练与大规模后训练升级基础模型Ling-2.0,其设计融合了模型架构、优化目标、服务系统与智能体训练环境的统一协同。架构层面引入混合线性注意力机制,结合Lightning Attention与MLA,显著提升长上下文训练与解码效率。为增强每输出词元的能力,采用进化链式思维、语言单元策略优化、双向偏好对齐及最短正确响应蒸馏等方法。针对智能体能力,提出KPop强化学习框架,通过异步调度编码、搜索、工具使用与工作流执行,支持在大规模环境相关数据上稳定训练Ring-2.6-1T。Ling-2.6与Ring-2.6共同构建了高效、可扩展且开放的智能体系统路径。所有2.6系列检查点已开源,以支持实用智能体智能的进一步研究与发展。

原文摘要 · Abstract (English)

Efficient and scalable agentic intelligence requires models that can deliver both low-latency responses and strong reasoning capabilities while remaining practical to train, serve, and deploy. In this report, we present Ling-2.6 and Ring-2.6, a family of models designed to address this challenge at scale. Ling-2.6 is optimized for instant response generation and high capability per output token, whereas Ring-2.6 is tailored for deeper reasoning and more advanced agentic workflows. Instead of training from scratch, we upgrade the Ling-2.0 base model through architectural migration pre-training and large-scale post-training. This upgrade is guided by a unified co-design of model architecture, optimization objectives, serving systems, and agent training environments, enabling improvements in both model capability and deployment efficiency. At the architectural level, we introduce a hybrid linear attention design that integrates Lightning Attention with MLA, improving the efficiency of long-context training and decoding. To further enhance token efficiency, we optimize capability per output token through Evolutionary Chain-of-Thought, Linguistic Unit Policy Optimization, bidirectional preference alignment, and shortest-correct-response distillation. For agentic capabilities, we propose KPop, a reinforcement learning framework designed to support stable training of Ring-2.6-1T on large-scale environment-grounded data. KPop improves training efficiency through asynchronous scheduling across coding, search, tool use, and workflow execution, enabling scalable learning from complex agent-environment interactions. Together, Ling-2.6 and Ring-2.6 provide a practical pathway toward efficient, scalable, and open agentic systems. We open-source all checkpoints in the 2.6 family to support further research and development in practical agentic intelligence.

智能体大模型推理优化开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。