arXiv:2507.20534cs.LGcs.AI2025-07被引 371

Kimi K2是开源的智能代理大模型,具备强大代码与推理能力。

Kimi K2: Open Agentic Intelligence

  • 采用专家混合架构,激活参数320亿,总参数达1万亿。
  • 训练零崩溃,多项评测超越多数开源与闭源模型。
  • 适合研究智能代理、软件工程及无需思考推理的任务。

我们推出Kimi K2,一款拥有320亿激活参数和1万亿总参数的专家混合(MoE)大语言模型。提出MuonClip优化器,通过QK-clip技术解决训练不稳定性,同时保持Muon的高效令牌处理能力。基于MuonClip,K2在15.5万亿令牌上完成预训练,无损失突增。后训练阶段包含大规模智能体数据合成流程与联合强化学习阶段,模型通过与真实和合成环境交互提升能力。Kimi K2在开源非思考模型中表现领先,尤其在智能体任务上:在Tau2-Bench得66.1分,ACEBench(En)76.5分,SWE-Bench Verified 65.8分,SWE-Bench Multilingual 47.3分,均优于多数开源与闭源基线。在代码、数学与推理任务中表现优异:LiveCodeBench v6得53.7分,AIME 2025 49.5分,GPQA-Diamond 75.1分,OJBench 27.1分,均未启用扩展思考。该模型为当前最强大的开源大模型之一,尤其适用于软件工程与智能体任务。我们公开基础模型与后训练检查点,推动智能体研究与应用。

原文摘要 · Abstract (English)

We introduce Kimi K2, a Mixture-of-Experts (MoE) large language model with 32 billion activated parameters and 1 trillion total parameters. We propose the MuonClip optimizer, which improves upon Muon with a novel QK-clip technique to address training instability while enjoying the advanced token efficiency of Muon. Based on MuonClip, K2 was pre-trained on 15.5 trillion tokens with zero loss spike. During post-training, K2 undergoes a multi-stage post-training process, highlighted by a large-scale agentic data synthesis pipeline and a joint reinforcement learning (RL) stage, where the model improves its capabilities through interactions with real and synthetic environments. Kimi K2 achieves state-of-the-art performance among open-source non-thinking models, with strengths in agentic capabilities. Notably, K2 obtains 66.1 on Tau2-Bench, 76.5 on ACEBench (En), 65.8 on SWE-Bench Verified, and 47.3 on SWE-Bench Multilingual -- surpassing most open and closed-sourced baselines in non-thinking settings. It also exhibits strong capabilities in coding, mathematics, and reasoning tasks, with a score of 53.7 on LiveCodeBench v6, 49.5 on AIME 2025, 75.1 on GPQA-Diamond, and 27.1 on OJBench, all without extended thinking. These results position Kimi K2 as one of the most capable open-source large language models to date, particularly in software engineering and agentic tasks. We release our base and post-trained model checkpoints to facilitate future research and applications of agentic intelligence.

智能代理代码生成大模型开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。