构建专家级多领域视频理解评测基准,推动模型从看图到懂专业。
MMVU: Measuring Expert-Level Multi-Discipline Video Understanding

- 基于27个学科3000道专家标注题,覆盖科学、医疗、人文、工程四领域。
- 引入专家推理链与领域知识,模型仅在o1和Gemini 2.0 Flash Thinking中表现最优。
- 适合研究多模态模型在专业知识场景下的推理能力提升方向。
我们提出MMVU,一个面向基础模型视频理解的专家级、跨学科评估基准。该基准包含3,000道由专家从头标注的问题,涵盖科学、医疗、人文与社会科学、工程四大核心领域中的27个学科。相较于现有基准,MMVU具有三大创新:首先,要求模型应用领域专属知识并进行专家级推理以分析专业领域视频,突破传统视频评测中对基础视觉感知的局限;其次,所有样本均由对应领域专家重新标注,并通过严格的数据质量控制保障数据可靠性;最后,每个样本均配有专家标注的推理过程与相关领域知识,支持深度分析。我们在MMVU上对32个前沿多模态基础模型进行了全面评估,具备System-2能力的最新模型o1与Gemini 2.0 Flash Thinking表现最佳,但仍显著落后于人类专家。通过深入错误分析与案例研究,我们为未来在特定领域实现知识密集型专家级视频理解提供了可操作的改进方向。
原文摘要 · Abstract (English)
We introduce MMVU, a comprehensive expert-level, multi-discipline benchmark for evaluating foundation models in video understanding. MMVU includes 3,000 expert-annotated questions spanning 27 subjects across four core disciplines: Science, Healthcare, Humanities & Social Sciences, and Engineering. Compared to prior benchmarks, MMVU features three key advancements. First, it challenges models to apply domain-specific knowledge and perform expert-level reasoning to analyze specialized-domain videos, moving beyond the basic visual perception typically assessed in current video benchmarks. Second, each example is annotated by human experts from scratch. We implement strict data quality controls to ensure the high quality of the dataset. Finally, each example is enriched with expert-annotated reasoning rationals and relevant domain knowledge, facilitating in-depth analysis. We conduct an extensive evaluation of 32 frontier multimodal foundation models on MMVU. The latest System-2-capable models, o1 and Gemini 2.0 Flash Thinking, achieve the highest performance among the tested models. However, they still fall short of matching human expertise. Through in-depth error analyses and case studies, we offer actionable insights for future advancements in expert-level, knowledge-intensive video understanding for specialized domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。