2.8万亿参数开源大模型,支持百万词上下文与视觉能力
Kimi K3: Open Frontier Intelligence
- 采用混合专家架构,每令牌激活16个专家,提升推理效率
- 长序列任务表现提升2.5倍,支持百万词上下文与复杂推理
- 适合研究前沿智能、长程任务执行与开源模型生态建设
我们介绍Kimi K3,一个拥有2.8万亿参数的混合专家(MoE)模型,具备1040亿激活参数、原生视觉能力以及100万词上下文窗口。该模型基于Kimi Delta Attention和注意力残差机制,优化了长序列与深层网络中的信息流动。结合稳定潜空间混合专家(Stable LatentMoE),每令牌仅激活896个路由专家中的16个,配合优化的训练与数据策略,使整体扩展效率相比Kimi K2提升约2.5倍。后训练阶段在通用、自主代理及编码领域引入强化学习,支持多层级推理,实现组合泛化与稳健的长周期执行。在2.8万亿规模下,得益于算法-系统协同设计(如KDA)、专家并行训练的完全平衡与高效内存管理、持久化回滚与沙盒状态的百万词自主强化学习,以及部署创新,评估显示其在长周期编码、自主代理、知识、推理和视觉任务中达到前沿水平。尽管整体性能仍略逊于顶级闭源模型(如Claude Fable 5和GPT-5.6 Sol),但其在评测套件中持续优于其他开源与闭源模型。我们开放全部模型权重,以促进未来研究并加速前沿智能的广泛部署。
原文摘要 · Abstract (English)
We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes, these advances yield an approximately 2.5x improvement in overall scaling efficiency over Kimi K2. Post-training highlights reinforcement learning across general, agentic, and coding domains and multiple reasoning-effort levels, enabling compositional generalization and robust long-horizon execution. At 2.8T scale, Kimi K3 is supported by infrastructure advances in multiple areas: algorithm-system co-design for KDA, perfectly balanced expert-parallel training with efficient memory management, million-token agentic RL with persistent rollout and sandbox states, and deployment innovations. Extensive evaluations show that Kimi K3 achieves frontier-level performance across long-horizon coding, agentic, knowledge, reasoning, and vision tasks. While its overall performance still trails the most powerful proprietary models, namely Claude Fable 5 and GPT-5.6 Sol, Kimi K3 consistently outperforms other open and proprietary models evaluated in our suite. We release the full Kimi K3 model weights to facilitate future research and accelerate the broader deployment and adoption of frontier intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。