arXiv:2501.00584cs.CVcs.LG2025-01CVPR被引 49

提出在线视频理解新框架,突破实时处理瓶颈。

Online Video Understanding: OVBench and VideoChat-Online

论文配图:Online Video Understanding: OVBench and VideoChat-Online
图 1 · 摘自论文原文
  • 设计金字塔记忆库,高效保留视频时空关键信息
  • 在OVBench上超越现有模型23.7%(在线)和4.19%(离线)
  • 适合自动驾驶、人机交互等实时视频场景应用

多模态大语言模型在离线视频理解方面已取得显著进展。然而,在自动驾驶、人机交互等真实场景中,面对连续的在线视频流,实时处理带来独特挑战。本文从评估基准、模型架构和训练策略三方面开展系统性工作:首先,提出OVBench,一个涵盖过去、当前、未来三类时间上下文的问答基准,包含6大任务类型与16个子任务,来自多个数据集;其次,提出金字塔记忆库(PMB),有效保留视频流中的关键时空信息;第三,设计离线到在线学习范式,构建针对在线视频的交错对话格式与指令微调数据集。基于此框架,开发出VideoChat-Online模型。尽管计算成本更低、效率更高,其在主流离线视频基准及OVBench上均优于现有先进离线模型Qwen2-VL 7B和在线模型Flash-VStream,分别提升4.19%和23.7%。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have significantly progressed in offline video understanding. However, applying these models to real-world scenarios, such as autonomous driving and human-computer interaction, presents unique challenges due to the need for real-time processing of continuous online video streams. To this end, this paper presents systematic efforts from three perspectives: evaluation benchmark, model architecture, and training strategy. First, we introduce OVBench, a comprehensive question-answering benchmark designed to evaluate models' ability to perceive, memorize, and reason within online video contexts. It features 6 core task types across three temporal contexts-past, current, and future-forming 16 subtasks from diverse datasets. Second, we propose a new Pyramid Memory Bank (PMB) that effectively retains key spatiotemporal information in video streams. Third, we proposed an offline-to-online learning paradigm, designing an interleaved dialogue format for online video data and constructing an instruction-tuning dataset tailored for online video training. This framework led to the development of VideoChat-Online, a robust and efficient model for online video understanding. Despite the lower computational cost and higher efficiency, VideoChat-Online outperforms existing state-of-the-art offline and online models across popular offline video benchmarks and OVBench, demonstrating the effectiveness of our model architecture and training strategy. % Our approach surpasses existing state-of-the-art offline models Qwen2-VL 7B and online models Flash-VStream, by 4.19% and 23.7% on OVBench, respectively.

视频理解在线推理大模型记忆机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。