Streamo让视频大模型实时互动,能懂动作、答问题、说场景。
Streaming Video Instruction Tuning
- 用46.5万条指令数据训练,支持多任务实时视频理解
- 在多个流式视频任务上表现超越现有模型,响应快且准确
- 适合需要实时视频分析的智能助手、机器人等场景
我们提出 Streamo,一种面向实时流媒体视频的通用交互式语言模型。与现有仅聚焦问答或字幕生成的在线视频模型不同,Streamo 能完成实时叙述、动作理解、事件字幕、时间定位和时敏问答等多种任务。为实现这一多样性,我们构建了 Streamo-Instruct-465K 数据集,覆盖多样时序上下文与多任务监督,支持跨异构任务的统一训练。通过简化流程端到端训练后,Streamo 展现出强大的时序推理能力、快速响应和广泛泛化性,在多种流式视频基准测试中表现优异。实验表明,Streamo 缩小了离线视频感知模型与实时多模态助手之间的差距,朝着统一的连续视频智能理解迈进一步。
原文摘要 · Abstract (English)
We present Streamo, a real-time streaming video LLM that serves as a general-purpose interactive assistant. Unlike existing online video models that focus narrowly on question answering or captioning, Streamo performs a broad spectrum of streaming video tasks, including real-time narration, action understanding, event captioning, temporal event grounding, and time-sensitive question answering. To develop such versatility, we construct Streamo-Instruct-465K, a large-scale instruction-following dataset tailored for streaming video understanding. The dataset covers diverse temporal contexts and multi-task supervision, enabling unified training across heterogeneous streaming tasks. After training end-to-end on the instruction-following dataset through a streamlined pipeline, Streamo exhibits strong temporal reasoning, responsive interaction, and broad generalization across a variety of streaming benchmarks. Extensive experiments show that Streamo bridges the gap between offline video perception models and real-time multimodal assistants, making a step toward unified, intelligent video understanding in continuous video streams.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。