arXiv:2511.07250cs.CVcs.AI2025-11NeurIPS被引 8

首个评估多视频理解能力的基准,填补了MMLM评测空白。

MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs

  • 构建涵盖8项核心能力的多视频评测集
  • 包含4959个视频与1824个问答对
  • 适合自动驾驶、体育分析等场景研究者

多模态大语言模型(MLLMs)已扩展至视觉领域,但现有评测基准仍局限于单视频理解,难以满足真实场景(如体育分析、自动驾驶)中多视频理解的需求。为此,我们提出首个面向MLLM多视频理解的综合性评测基准MVU-Eval。该基准通过4,959个来自不同领域的视频,构建了1,824个精心设计的问答对,评估八项核心能力,涵盖基础感知与高阶推理任务,与多传感器融合、跨视角体育分析等实际应用紧密对齐。对主流开源与闭源模型的全面评估揭示了当前MLLM在多视频理解上的显著性能差距与局限性。该基准将公开发布,以推动后续研究。

原文摘要 · Abstract (English)

The advent of Multimodal Large Language Models (MLLMs) has expanded AI capabilities to visual modalities, yet existing evaluation benchmarks remain limited to single-video understanding, overlooking the critical need for multi-video understanding in real-world scenarios (e.g., sports analytics and autonomous driving). To address this significant gap, we introduce MVU-Eval, the first comprehensive benchmark for evaluating Multi-Video Understanding for MLLMs. Specifically, our MVU-Eval mainly assesses eight core competencies through 1,824 meticulously curated question-answer pairs spanning 4,959 videos from diverse domains, addressing both fundamental perception tasks and high-order reasoning tasks. These capabilities are rigorously aligned with real-world applications such as multi-sensor synthesis in autonomous systems and cross-angle sports analytics. Through extensive evaluation of state-of-the-art open-source and closed-source models, we reveal significant performance discrepancies and limitations in current MLLMs' ability to perform understanding across multiple videos. The benchmark will be made publicly available to foster future research.

多视频理解评测基准MLLM自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。