arXiv:2507.01949cs.CV2025-07被引 32

80亿参数模型专攻短视频理解,推理能力更强。

Kwai Keye-VL Technical Report

  • 四阶段预训练+双阶段后训练,强化视觉语言对齐
  • 在视频基准上达顶尖表现,图像任务也保持竞争力
  • 新基准KC-MMBench验证真实场景优势,适合短视频研究

尽管多模态大语言模型在静态图像上表现出色,但在动态、信息密集的短视频理解方面仍显不足,而短视频是当今数字内容的主要形式。为此,我们推出80亿参数的多模态基础模型Kwai Keye-VL,专为领先短视频理解能力设计,同时保持强大的通用视觉-语言能力。其开发基于两大支柱:一个超过6000亿标记、以视频为重点的高质量大规模数据集,以及创新的训练方法。该方法包含四阶段预训练实现稳固的视觉-语言对齐,随后是两阶段后训练:第一阶段提升指令遵循等基础能力,第二阶段聚焦激发高级推理能力。关键创新在于五模式‘冷启动’数据混合,包含‘思考’、‘非思考’、‘自动思考’、‘带图像思考’及高质量视频数据,教会模型自主判断何时何地进行推理。后续强化学习与对齐步骤进一步优化推理能力,并纠正重复输出等异常行为。大量评估显示,Keye-VL在公开视频基准上达到最先进水平,在通用图像任务中仍具竞争力(图1)。此外,我们构建并发布专为真实短视频场景设计的基准KC-MMBench,Keye-VL在此展现显著优势。

原文摘要 · Abstract (English)

While Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities on static images, they often fall short in comprehending dynamic, information-dense short-form videos, a dominant medium in today's digital landscape. To bridge this gap, we introduce \textbf{Kwai Keye-VL}, an 8-billion-parameter multimodal foundation model engineered for leading-edge performance in short-video understanding while maintaining robust general-purpose vision-language abilities. The development of Keye-VL rests on two core pillars: a massive, high-quality dataset exceeding 600 billion tokens with a strong emphasis on video, and an innovative training recipe. This recipe features a four-stage pre-training process for solid vision-language alignment, followed by a meticulous two-phase post-training process. The first post-training stage enhances foundational capabilities like instruction following, while the second phase focuses on stimulating advanced reasoning. In this second phase, a key innovation is our five-mode ``cold-start'' data mixture, which includes ``thinking'', ``non-thinking'', ``auto-think'', ``think with image'', and high-quality video data. This mixture teaches the model to decide when and how to reason. Subsequent reinforcement learning (RL) and alignment steps further enhance these reasoning capabilities and correct abnormal model behaviors, such as repetitive outputs. To validate our approach, we conduct extensive evaluations, showing that Keye-VL achieves state-of-the-art results on public video benchmarks and remains highly competitive on general image-based tasks (Figure 1). Furthermore, we develop and release the \textbf{KC-MMBench}, a new benchmark tailored for real-world short-video scenarios, where Keye-VL shows a significant advantage.

视频理解多模态模型推理能力短视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。