让视觉语言模型实时处理真实世界视频流,支持长时间记忆与快速响应。
Harnessing Streaming Video in the Wild

- 构建新数据集和训练目标,让模型适应持续视频流交互。
- 推出可插拔系统,实现秒级响应、12小时记忆、亚秒延迟。
- 发布评估基准,推动社区从离线理解迈向实际部署的流式智能。
视觉语言模型(VLMs)在视频通话助手、直播解说和具身机器人等场景中越来越需要处理无边界视频流。理想的流式系统应支持主动交互、长时记忆和实时处理,且基于能应对多样化真实场景的VLM骨干网络。然而现有VLM在离线视频理解上表现优异,但在流式能力上不足,缺乏专用部署基础设施。本文从三方面填补该空白:(i) 构建Streaming-Train-248K数据集及新型训练目标,提升VLM对流式交互与理解的能力;(ii) 提出Streaming Harness系统,为任意VLM赋予主动交互(每秒响应决策)、长时记忆(12小时上下文保留)和实时处理(亚秒级延迟)三大核心能力;(iii) 设计Streaming-Eval基准,评估模型在多种真实场景下的流式能力。大量实验表明,本方法在所有关键能力上均有稳定提升。我们将开源数据、代码与基准,推动社区从离线视频理解向可部署的流式智能转型。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) are increasingly required to process unbounded video streams in applications such as video-call assistants, live commentary, and embodied robots. An ideal streaming system should support proactive interaction, long-horizon memory, and real-time processing, while resting on a VLM backbone capable of handling diverse in-the-wild streaming tasks. However, existing VLMs excel at offline video understanding but fall short in streaming capabilities and lack dedicated infrastructure for streaming deployment. We address this gap on three fronts. (i) For backbone capability, we construct \textbf{Streaming-Train-248K}, a streaming dataset paired with a novel training objective for adapting VLMs to streaming interaction and understanding. (ii) For real-world deployment, we introduce \textbf{Streaming Harness}, a plug-and-play system that endows any VLM with three core abilities: proactive interaction (per-second response decisions), long-term memory (12-hour context retention), and real-time processing (sub-second latency). (iii) To drive continued community progress on streaming capabilities, we design \textbf{Streaming-Eval}, a benchmark that reflects models' capabilities across diverse in-the-wild scenarios. Extensive experiments demonstrate consistent gains from our approach across all core capabilities required for streaming video understanding. We will open-source our data, code, and benchmark to advance the community's shift from offline video understanding to deployable streaming intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。