arXiv:2606.10651cs.CV2026-06被引 1

30B参数的开源多模态模型,支持256K上下文长视频理解与智能体协作。

Kwai Keye-VL-2.0 Technical Report

论文配图:Kwai Keye-VL-2.0 Technical Report
图 1 · 摘自论文原文
  • 采用稀疏注意力与专家混合架构,实现无损256K长视频处理。
  • 激活仅30亿参数,在多个长视频基准上达顶尖性能。
  • 支持代码、工具、搜索等场景的多模态自主纠错与协同智能。

我们介绍 Kwai Keye-VL-2.0-30B-A3B,一个开源的专家混合(MoE)多模态基础模型,旨在推进长视频理解与智能体智能。为解决小时级视频中的超长上下文、信息冗余和计算成本过高问题,Keye-VL-2.0首次将 DeepSeek 稀疏注意力(DSA)适配至基于分组查询注意力(GQA)的多模态架构,实现无损256K上下文处理,同时捕捉关键帧与长时程依赖。该架构依托高度优化的训练与推理基础设施,包括可扩展视频输入输出、异构ViT-LM并行机制及自定义DSA内核,显著提升吞吐量并降低计算开销。为克服多任务对齐中灾难性遗忘的算法困境,我们提出跨模态多教师在线蒸馏(MOPD)结合上下文强化学习(Context-RL)与视频强化学习(Video-RL)。通过将在线回放生成的密集令牌级教师反馈蒸馏回仅激活30亿参数的MoE主干网络,Keye-VL-2.0原生支持在代码、工具、搜索场景中实现多模态自我修正的高级智能体协作。在视频理解、时间定位、推理、STEM及智能体基准上的广泛评估表明,Keye-VL-2.0-30B-A3B在同类规模模型中达到最先进水平,尤其在TimeLens的细粒度时间定位,以及Video-MME-v2和LongVideoBench的长视频理解任务中表现突出。我们发布模型检查点,以加速社区在可扩展、鲁棒的多模态智能体应用方面的进展。

原文摘要 · Abstract (English)

We introduce Kwai Keye-VL-2.0-30B-A3B, an open-source Mixture-of-Experts (MoE) multimodal foundation model designed to advance long-video understanding and agentic intelligence. To address the challenges of ultra-long contexts, information redundancy, and prohibitive computational costs inherent in hour-level videos, Keye-VL-2.0 is the first to adapt DeepSeek Sparse Attention (DSA) to GQA-based multimodal architectures, enabling lossless 256K context processing while capturing critical frames and long-range temporal dependencies. This architecture is underpinned by a highly optimized training and inference infrastructure, including scalable video I/O, heterogeneous ViT-LM parallelism, and custom DSA kernels that significantly maximize throughput and minimize computational overhead. Furthermore, to overcome the algorithmic dilemma of catastrophic forgetting during multi-task alignment, we introduce Cross-Modal Multi-Teacher On-Policy Distillation (MOPD) paired with Context-RL and Video-RL. By distilling dense token-level teacher feedback from on-policy rollouts back into the MoE backbone, which activates only 3B parameters, Keye-VL-2.0 natively empowers advanced agent collaboration across Code, Tool, and Search scenarios with multimodal self-correction. Extensive evaluations across video understanding, temporal grounding, reasoning, STEM, and agent benchmarks demonstrate that Keye-VL-2.0-30B-A3B achieves state-of-the-art performance among models of similar scale, particularly excelling in fine-grained temporal localization on TimeLens and long-video comprehension on Video-MME-v2 and LongVideoBench. We release our model checkpoints to accelerate community progress toward scalable and robust multimodal agentic applications.

多模态长视频智能体MoE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。