arXiv:2501.13826cs.CVcs.CL2025-01被引 228

评测大模型从专业视频中获取知识的能力,发现模型远不如人类。

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

  • 设计六学科视频评测集,分感知、理解、应用三阶段提问。
  • 引入知识增益指标Δknowledge,量化看视频后的能力提升。
  • 模型在高阶任务表现差,揭示当前大模型视频学习短板。

人类通过感知、理解、应用三个认知阶段获取知识,视频是促进这一过程的有效媒介。然而现有视频基准无法系统评估大视听模型(LMMs)的知识获取能力。为此,我们提出Video-MMMU,一个跨多模态、多学科的评测基准,包含300个专家级视频和900个由人工标注的问题,覆盖六个学科领域。通过与认知阶段对齐的问答对,评估模型在感知、理解、应用三个层面的知识获取能力。我们提出知识增益指标Δknowledge,用于量化观看视频后性能的提升。对LMMs的评估显示,随着认知需求增加,模型性能急剧下降,且与人类存在显著差距,凸显了提升大模型从视频中学习和适应能力的迫切需求。

原文摘要 · Abstract (English)

Humans acquire knowledge through three cognitive stages: perceiving information, comprehending knowledge, and adapting knowledge to solve novel problems. Videos serve as an effective medium for this learning process, facilitating a progression through these cognitive stages. However, existing video benchmarks fail to systematically evaluate the knowledge acquisition capabilities in Large Multimodal Models (LMMs). To address this gap, we introduce Video-MMMU, a multi-modal, multi-disciplinary benchmark designed to assess LMMs' ability to acquire and utilize knowledge from videos. Video-MMMU features a curated collection of 300 expert-level videos and 900 human-annotated questions across six disciplines, evaluating knowledge acquisition through stage-aligned question-answer pairs: Perception, Comprehension, and Adaptation. A proposed knowledge gain metric, Δknowledge, quantifies improvement in performance after video viewing. Evaluation of LMMs reveals a steep decline in performance as cognitive demands increase and highlights a significant gap between human and model knowledge acquisition, underscoring the need for methods to enhance LMMs' capability to learn and adapt from videos.

视频理解多模态知识获取评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。