arXiv:2504.14693cs.CVcs.AI2025-04ICCV被引 43

构建视频讲座理解大基准,测试模型跨学科认知能力

Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark

  • 设计大规模多学科讲座评测集,覆盖多种教学场景
  • 90余款模型评估显示,现有模型在感知与推理结合任务上表现有限
  • 揭示视觉帧数与大模型规模对讲座理解的影响规律

语言多模态模型(LMMs)在视频理解方面取得进展,但多学科讲座理解仍待探索。我们提出 Video-MMLU,一个用于评估 LMMs 理解多学科讲座能力的大规模基准。我们评测了超过 90 款开源与专有模型,参数量从 0.5B 到 40B。结果表明,当前模型在需结合感知与推理的任务中存在明显局限。此外,我们研究了视觉帧数与大语言模型规模对性能的影响,揭示了多模态感知与推理在讲座理解中的交互机制。

原文摘要 · Abstract (English)

Recent advancements in language multimodal models (LMMs) for video have demonstrated their potential for understanding video content, yet the task of comprehending multi-discipline lectures remains largely unexplored. We introduce Video-MMLU, a massive benchmark designed to evaluate the capabilities of LMMs in understanding Multi-Discipline Lectures. We evaluate over 90 open-source and proprietary models, ranging from 0.5B to 40B parameters. Our results highlight the limitations of current models in addressing the cognitive challenges presented by these lectures, especially in tasks requiring both perception and reasoning. Additionally, we explore how the number of visual tokens and the large language models influence performance, offering insights into the interplay between multimodal perception and reasoning in lecture comprehension.

多模态理解视频评测讲座理解大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。