arXiv:2411.10979cs.CVcs.AI2024-11CVPR被引 23

评测大模型对影视级视频构图的理解能力,发现模型表现远逊于人类。

VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?

  • 用精心编排的影视级视频和专业标注构建新评估基准
  • 33个模型在982段视频上平均准确率仅达42.3%
  • 适合研究视频理解、多模态模型评估的学者与工程师

多模态大语言模型(MLLMs)在多模态理解方面取得显著进展,拓展了视频内容分析能力。然而,现有评估基准主要聚焦于抽象视频理解,缺乏对视频构图能力的细致评估——即视觉元素在高度剪辑视频中如何组合与互动的深层理解。我们提出VidComposition,一个专门用于评估MLLM视频构图理解能力的新基准,包含982段视频和1706道多选题,覆盖镜头运动、角度、景别、叙事结构、角色动作与情绪等构图维度。对33个开源与专有MLLM的全面评估显示,模型性能与人类存在显著差距。这揭示了当前MLLM在复杂剪辑视频构图理解上的局限性,并为未来改进提供了方向。排行榜与评估代码已公开:https://yunlong10.github.io/VidComposition/

原文摘要 · Abstract (English)

The advancement of Multimodal Large Language Models (MLLMs) has enabled significant progress in multimodal understanding, expanding their capacity to analyze video content. However, existing evaluation benchmarks for MLLMs primarily focus on abstract video comprehension, lacking a detailed assessment of their ability to understand video compositions, the nuanced interpretation of how visual elements combine and interact within highly compiled video contexts. We introduce VidComposition, a new benchmark specifically designed to evaluate the video composition understanding capabilities of MLLMs using carefully curated compiled videos and cinematic-level annotations. VidComposition includes 982 videos with 1706 multiple-choice questions, covering various compositional aspects such as camera movement, angle, shot size, narrative structure, character actions and emotions, etc. Our comprehensive evaluation of 33 open-source and proprietary MLLMs reveals a significant performance gap between human and model capabilities. This highlights the limitations of current MLLMs in understanding complex, compiled video compositions and offers insights into areas for further improvement. The leaderboard and evaluation code are available at https://yunlong10.github.io/VidComposition/

视频理解多模态评估基准构图分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。