评测大模型在实时视频流中的感知、理解与推理能力。
RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video

- 构建多时戳问答与分层问题结构,细粒度评估连续视频理解。
- 覆盖552段视频、4608个问答对,发现实时模型性能优于离线模型。
- 揭示当前架构处理长视频流的局限性,适合关注视频理解的研究者。
多模态大语言模型在感知、理解与推理方面进展迅速,但现有基准难以评估其在连续动态真实视频流中的表现。此类场景要求模型在视觉场景随时间演变时保持连贯的理解与推理能力。本文提出RTV-Bench,一个面向多模态大语言模型实时视频分析的细粒度基准。该基准基于三大原则:多时戳问答、跨感知与推理的分层问题结构,以及对连续感知、理解与推理的多维评估。RTV-Bench包含552段多样化视频和4,608个精心设计的问答对,覆盖广泛动态场景。我们评估了多种前沿多模态大模型,包括专有模型、开源离线模型及开源实时模型。结果表明,实时模型总体优于离线模型,但仍落后于顶尖专有系统。模型容量扩大通常带来性能提升,但单纯增加采样帧密度并未一致提高效果。这表明当前架构在处理长时视频流时存在固有局限,亟需为流式视频处理专门设计的模型。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have made rapid progress in perception, understanding, and reasoning, yet existing benchmarks fall short in evaluating these abilities under continuous and dynamic real-world video streams. Such settings require models to maintain coherent understanding and reasoning as visual scenes evolve over time. **We introduce RTV-Bench, a fine-grained benchmark for real-time video analysis with MLLMs**. It is built upon three key principles: multi-timestamp question answering, hierarchical question structures spanning perception and reasoning, and multi-dimensional evaluation of continuous perception, understanding, and reasoning. RTV-Bench comprises 552 diverse videos and 4,608 carefully curated QA pairs covering a wide range of dynamic scenarios. We evaluate a broad range of state-of-the-art MLLMs, including proprietary, open-source offline, and open-source real-time models. Our results show that real-time models generally outperform offline counterparts but still lag behind leading proprietary systems. While scaling model capacity generally yields performance gains, simply increasing the density of sampled input frames does not consistently translate into improved results. These observations suggest inherent limitations in current architectures when handling long-horizon video streams, underscoring the need for models explicitly designed for streaming video processing and analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。