Kimi K2.5通过多模态协同与智能体集群,实现高效复杂任务处理。
Kimi K2.5: Visual Agentic Intelligence

- 融合文本与视觉的联合优化,提升多模态理解能力。
- 在编码、视觉、推理等任务上达到顶尖水平,延迟降低4.5倍。
- 适合研究智能体系统与多模态应用的开发者与学者。
我们提出Kimi K2.5,一个开源的多模态智能体模型,旨在推动通用智能体智能的发展。该模型强调文本与视觉的联合优化,通过联合文本-视觉预训练、零视觉监督微调以及联合文本-视觉强化学习等技术,使两种模态相互增强。在此多模态基础上,Kimi K2.5引入Agent Swarm——一种自驱动的并行智能体编排框架,可动态将复杂任务分解为异构子问题并并发执行。大量评估显示,Kimi K2.5在编码、视觉、推理和智能体任务等多个领域均取得当前最优性能,且相比单智能体基线,延迟降低高达4.5倍。我们已公开后训练的Kimi K2.5模型检查点,以促进未来智能体智能的研究与实际应用。
原文摘要 · Abstract (English)
We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that two modalities enhance each other. This includes a series of techniques such as joint text-vision pre-training, zero-vision SFT, and joint text-vision reinforcement learning. Building on this multimodal foundation, K2.5 introduces Agent Swarm, a self-directed parallel agent orchestration framework that dynamically decomposes complex tasks into heterogeneous sub-problems and executes them concurrently. Extensive evaluations show that Kimi K2.5 achieves state-of-the-art results across various domains including coding, vision, reasoning, and agentic tasks. Agent Swarm also reduces latency by up to $4.5\times$ over single-agent baselines. We release the post-trained Kimi K2.5 model checkpoint to facilitate future research and real-world applications of agentic intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。