首个评估大模型流式视频理解能力的基准,揭示当前模型与人类差距。
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
- 构建18项任务、900个视频的流式视频评测集,模拟实时问答场景。
- 13个主流模型在测试中表现远低于人类水平,顶尖模型如Gemini 1.5 Pro也明显不足。
- 适合研究多模态大模型实时推理、视频理解与人机交互的学者和开发者。
多模态大语言模型(MLLMs)已从图像理解扩展到视频理解,但多数仍聚焦于离线处理,需完整读取所有视频帧才能响应查询,无法实现真正的实时流式理解。本文提出StreamingBench,首个全面评估MLLM流式视频理解能力的基准。该基准涵盖三个核心维度:实时视觉理解、多源信息理解与上下文理解,包含18个任务、900个视频及4,500组人工标注问答对。每段视频在不同时间点设置5个问题,模拟连续流式输入。在13个开源与专有模型上进行实验发现,即使最先进的模型如Gemini 1.5 Pro和GPT-4o,在流式视频理解方面仍显著落后于人类水平。本工作旨在推动MLLM向真实场景下的人类级视频感知与交互迈进。
原文摘要 · Abstract (English)
The rapid development of Multimodal Large Language Models (MLLMs) has expanded their capabilities from image comprehension to video understanding. However, most of these MLLMs focus primarily on offline video comprehension, necessitating extensive processing of all video frames before any queries can be made. This presents a significant gap compared to the human ability to watch, listen, think, and respond to streaming inputs in real time, highlighting the limitations of current MLLMs. In this paper, we introduce StreamingBench, the first comprehensive benchmark designed to evaluate the streaming video understanding capabilities of MLLMs. StreamingBench assesses three core aspects of streaming video understanding: (1) real-time visual understanding, (2) omni-source understanding, and (3) contextual understanding. The benchmark consists of 18 tasks, featuring 900 videos and 4,500 human-curated QA pairs. Each video features five questions presented at different time points to simulate a continuous streaming scenario. We conduct experiments on StreamingBench with 13 open-source and proprietary MLLMs and find that even the most advanced proprietary MLLMs like Gemini 1.5 Pro and GPT-4o perform significantly below human-level streaming video understanding capabilities. We hope our work can facilitate further advancements for MLLMs, empowering them to approach human-level video comprehension and interaction in more realistic scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。