arXiv:2601.02536cs.CV2026-01被引 4

构建首个无需参考答案的开放问答视频理解基准,评估多模态推理能力。

MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering Benchmark

  • 基于电影解说视频生成需跨模态推理的开放式问题
  • 8200+问题覆盖视觉、对话及两者结合,支持无参考评估
  • 发现视觉感知是模型瓶颈,且去视觉输入反而提升答案准确性

理解真实世界视频(如电影)需融合视觉与对话线索。现有VideoQA基准难以捕捉多模态推理,且因难以评估自由回答,多采用简单选择题。我们提出新型开放问答多模态VideoQA基准MovieRecapsQA,基于60个电影解说视频(YouTube内容)生成约8200个问题,配套所需答案事实信息;前者支持生成需多模态推理的问题,后者实现无需参考答案的评估机制。这是首个无参考开放问答VideoQA基准。每道题可在四种输入设置下评估:(a)完整电影,(b)完整解说视频(仅视觉),(c)约14分钟相关电影片段,(d)约1.2分钟对齐解说片段;所有情况均提供电影对话文本。问题按所需模态分类(视觉、对话或双模态),支持细粒度评估。我们在设置(d)上测试7个先进多模态大模型,发现:(i)仅本研究提出的无参考度量能产生与人类判断一致的模型排序;(ii)以视觉为核心的题目得分最低;(iii)移除视觉输入常能提升答案真实性;(iv)主要瓶颈为视觉感知而非视觉推理。

原文摘要 · Abstract (English)

Understanding real-world videos such as movies requires integrating visual and dialogue cues. Yet existing VideoQA benchmarks struggle to capture this multimodal reasoning and, given the difficulty of evaluating free-form answers, largely resort to simple multiple choice questions. We introduce a novel open-ended multimodal VideoQA benchmark, MovieRecapsQA, created using movie recap videos -- a distinctive type of YouTube content that summarizes a film via a voiceover description of key clips from the movie (recap video). From the transcribed voiceover (recap summary) of 60 recap videos, we generate $\approx$8.2K questions along with the necessary ``facts'' expected in each answer; the former facilitates the creation of questions that require mutimodal reasoning and the latter allow the construction of a reference-free evaluation metric that can be applied to open-ended responses. To our knowledge, this is the first reference-free open-ended VideoQA benchmark. The benchmark allows each question to be evaluated in different input video settings: given (a) the full-length movie, (b) the full ($\approx$11 min) recap video (visual only), (c) $\approx$14 min of aligned movie scenes, i.e, movie scenes relevant to the question, and (d) $\approx$1.2 min of aligned recap video scenes. In all cases, the text of any associated movie dialogue is provided. Each question is categorized by the modality required to answer it -- visual, dialogue, or both -- enabling fine-grained evaluation of multimodal capabilities. We benchmark (setting (d)) seven state-of-the-art MLLMs and find that (i) only our reference-free metric produces meaningful human-aligned model separation; (ii) vision-centric questions yield the lowest scores across all models; (iii) removing visual input often \textit{improves} model factuality; and (iv) the primary bottleneck is visual perception, not visual reasoning.

视频理解多模态开放问答评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。