构建首个多视频感知评估基准,测试模型跨视频理解能力。
MVPBench: A Multi-Video Perception Evaluation Benchmark for Multi-Modal Video Understanding
- 设计14项跨视频任务,覆盖多样视觉场景
- 包含5000道题、2700段视频,含人工标注数据
- 揭示现有模型在多视频理解上存在明显短板
大语言模型(LLMs)的快速发展推动了多模态大语言模型(MLLMs)的研究,并催生了评估其感知与理解能力的基准。然而,现有基准多局限于静态图像或单个视频,忽视了多个视频之间的复杂交互。为填补这一空白,我们提出多视频感知评估基准(MVPBench),包含14个子任务,覆盖多种视觉领域,旨在评估模型从视频序列中提取相关信息并做出决策的能力。MVPBench包含5000道问答题,涉及2700段来自现有数据集及人工标注的视频片段。大规模评估显示,当前模型在处理多视频输入时表现不佳,暴露出其在多视频理解上的显著局限。我们期待MVPBench能推动多视频感知技术的发展。
原文摘要 · Abstract (English)
The rapid progress of Large Language Models (LLMs) has spurred growing interest in Multi-modal LLMs (MLLMs) and motivated the development of benchmarks to evaluate their perceptual and comprehension abilities. Existing benchmarks, however, are limited to static images or single videos, overlooking the complex interactions across multiple videos. To address this gap, we introduce the Multi-Video Perception Evaluation Benchmark (MVPBench), a new benchmark featuring 14 subtasks across diverse visual domains designed to evaluate models on extracting relevant information from video sequences to make informed decisions. MVPBench includes 5K question-answering tests involving 2.7K video clips sourced from existing datasets and manually annotated clips. Extensive evaluations reveal that current models struggle to process multi-video inputs effectively, underscoring substantial limitations in their multi-video comprehension. We anticipate MVPBench will drive advancements in multi-video perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。