Keye-VL-1.5通过动态编码与长序列训练,大幅提升视频理解能力。
Kwai Keye-VL 1.5 Technical Report
- 采用慢快路径动态编码,按帧间变化分配计算资源。
- 支持最长128K token的上下文长度,可处理超长视频。
- 适合需要强视频推理与人类偏好对齐的应用场景。
近年来,大语言模型(LLMs)发展迅速,通过多模态大语言模型(MLLMs)拓展至多模态任务。然而,视频理解因内容动态且信息密集仍具挑战性,现有模型在空间分辨率与时间覆盖范围间存在权衡。本文提出Keye-VL-1.5,通过三项关键创新解决视频理解核心难题:首先,引入新型慢-快视频编码策略,根据帧间相似性动态分配计算资源,对视觉变化显著的关键帧以高分辨率处理(慢路径),对相对静态帧则以更高时间覆盖率低分辨率处理(快路径);其次,采用渐进式四阶段预训练方法,系统性将模型上下文长度从8K扩展至128K token,支持更长视频与复杂视觉内容处理;第三,构建全面后训练流程,聚焦推理增强与人类偏好对齐,包含五步链式思维数据构建、基于迭代GSPO的强化学习与渐进提示引导,以及对齐训练。在公开基准与严格的内部人工评估中,Keye-VL-1.5显著优于现有模型,尤其在视频理解任务上表现突出,同时在通用多模态基准上保持竞争力。
原文摘要 · Abstract (English)
In recent years, the development of Large Language Models (LLMs) has significantly advanced, extending their capabilities to multimodal tasks through Multimodal Large Language Models (MLLMs). However, video understanding remains a challenging area due to the dynamic and information-dense nature of videos. Existing models struggle with the trade-off between spatial resolution and temporal coverage when processing video content. We present Keye-VL-1.5, which addresses fundamental challenges in video comprehension through three key innovations. First, we introduce a novel Slow-Fast video encoding strategy that dynamically allocates computational resources based on inter-frame similarity, processing key frames with significant visual changes at higher resolution (Slow pathway) while handling relatively static frames with increased temporal coverage at lower resolution (Fast pathway). Second, we implement a progressive four-stage pre-training methodology that systematically extends the model's context length from 8K to 128K tokens, enabling processing of longer videos and more complex visual content. Third, we develop a comprehensive post-training pipeline focusing on reasoning enhancement and human preference alignment, incorporating a 5-step chain-of-thought data construction process, iterative GSPO-based reinforcement learning with progressive prompt hinting for difficult cases, and alignment training. Through extensive evaluation on public benchmarks and rigorous internal human assessment, Keye-VL-1.5 demonstrates significant improvements over existing models, particularly excelling in video understanding tasks while maintaining competitive performance on general multimodal benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。