提出新模型,让视频流处理效率提升80%且不丢关键信息。
TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos

- 用差分令牌过滤技术自动删掉视频中冗余帧。
- 实验显示删掉82.8%的视觉标记仍保持98%准确率。
- 适合做实时视频交互、直播分析等场景的应用者看。
在线视频平台,尤其是直播服务的快速发展,迫切需要能够实时理解视频流的系统。现有视频大模型虽擅长处理完整视频,但在连续流媒体场景下因无法高效处理密集冗余帧而受限。我们提出TimeChat-Online,一种新型在线视频大模型。其核心是差分令牌丢弃(DTD)模块,受人类视觉感知中的变化盲视现象启发,能保留有意义的时间变化,同时过滤帧间静态冗余内容。实验表明,DTD可实现82.8%的视频令牌减少,同时在StreamingBench上保持98%的性能,揭示流媒体视频中超过80%的视觉内容本质上冗余,无需语言引导即可去除。为支持无缝实时交互,我们构建了TimeChat-Online-139K数据集,涵盖回溯追踪、当前感知和未来响应等多种互动模式。该模型通过持续监测场景切换实现主动响应能力,优于传统方法。大量评估显示,它在StreamingBench和OvOBench上表现优异,并在长视频任务如Video-MME和MLVU上保持竞争力。
原文摘要 · Abstract (English)
The rapid growth of online video platforms, particularly live streaming services, has created an urgent need for real-time video understanding systems. These systems must process continuous video streams and respond to user queries instantaneously, presenting unique challenges for current Video Large Language Models (VideoLLMs). While existing VideoLLMs excel at processing complete videos, they face significant limitations in streaming scenarios due to their inability to handle dense, redundant frames efficiently. We introduce TimeChat-Online, a novel online VideoLLM that revolutionizes real-time video interaction. At its core lies our innovative Differential Token Drop (DTD) module, which addresses the fundamental challenge of visual redundancy in streaming videos. Drawing inspiration from human visual perception's Change Blindness phenomenon, DTD preserves meaningful temporal changes while filtering out static, redundant content between frames. Remarkably, our experiments demonstrate that DTD achieves an 82.8% reduction in video tokens while maintaining 98% performance on StreamingBench, revealing that over 80% of visual content in streaming videos is naturally redundant without requiring language guidance. To enable seamless real-time interaction, we present TimeChat-Online-139K, a comprehensive streaming video dataset featuring diverse interaction patterns including backward-tracing, current-perception, and future-responding scenarios. TimeChat-Online's unique Proactive Response capability, naturally achieved through continuous monitoring of video scene transitions via DTD, sets it apart from conventional approaches. Our extensive evaluation demonstrates TimeChat-Online's superior performance on streaming benchmarks (StreamingBench and OvOBench) and maintaining competitive results on long-form video tasks such as Video-MME and MLVU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。