arXiv:2506.06275cs.CVcs.CL2025-06被引 5

构建电影长视频理解新基准,测试模型对核心叙事的推理能力

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding

  • 设计真假陈述对,评估模型对剧情关键信息的理解
  • 850个陈述对中模型准确率远低于人类,凸显当前模型短板
  • 适合研究视频理解、叙事推理的学者与工程师使用

尽管视觉语言模型(VLMs)取得进展,但对长视频内容的整体理解仍具挑战性,部分原因在于现有基准存在局限。许多基准聚焦于‘大海捞针’式细节,鼓励模型进行上下文无关的检索,而非深度理解;另一些则依赖大规模半自动生成的问题(常由语言模型生成),虽易回答却无法反映真实理解力。本文提出MF$^2$,一个用于评估模型是否能理解、整合并回忆完整电影(时长50-170分钟)关键叙事信息的新基准。该数据集包含超过50部开源电影,每部配以人工构造的陈述对(一真一假),共超850对。这些陈述针对角色动机、情感、因果链和事件顺序等核心叙事元素,并指向人类无需重看即可回忆的关键场景。采用二元判断协议:每对中需正确识别真假陈述,避免选项顺序偏差,实现更精确的推理评估。实验表明,无论开源还是闭源的顶尖模型在该任务上均显著落后于人类,突显人类在保留与推理关键叙事信息上的优势,而当前VLM尚不具备此能力。

原文摘要 · Abstract (English)

Despite recent progress in vision-language models (VLMs), holistic understanding of long-form video content remains a significant challenge, partly due to limitations in current benchmarks. Many focus on peripheral, ``needle-in-a-haystack'' details, encouraging context-insensitive retrieval over deep comprehension. Others rely on large-scale, semi-automatically generated questions (often produced by language models themselves) that are easier for models to answer but fail to reflect genuine understanding. In this paper, we introduce MF$^2$, a new benchmark for evaluating whether models can comprehend, consolidate, and recall key narrative information from full-length movies (50-170 minutes long). MF$^2$ includes over 50 full-length, open-licensed movies, each paired with manually constructed sets of claim pairs -- one true (fact) and one plausible but false (fib), totalling over 850 pairs. These claims target core narrative elements such as character motivations and emotions, causal chains, and event order, and refer to memorable moments that humans can recall without rewatching the movie. Instead of multiple-choice formats, we adopt a binary claim evaluation protocol: for each pair, models must correctly identify both the true and false claims. This reduces biases like answer ordering and enables a more precise assessment of reasoning. Our experiments demonstrate that both open-weight and closed state-of-the-art models fall well short of human performance, underscoring the relative ease of the task for humans and their superior ability to retain and reason over critical narrative information -- an ability current VLMs lack.

视频理解叙事推理基准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。