arXiv:2506.21742cs.CV2025-06中稿 · CVPR被引 8

挑战视频理解隐含关系,测试模型推理能力

VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues

  • 从电影等创意视频中提取需推断的隐性关系问题
  • 1000个问答对,最佳模型仅64%准确率
  • 适合研究深度视频推理与人类认知对齐的学者

视频问答(VideoQA)通过多模态学习在对齐视觉与文本方面取得显著进展。然而,现有基准大多聚焦于可通过单帧或短片段直接观察到的动作、物体和事件回答的问题。要真正像人类一样理解视频,模型必须超越显性内容,推断跨帧隐含的关系与上下文线索。当前基准未能捕捉这一核心能力。为此,我们提出VRR-QA,一个面向显性线索之外视觉关系推理的新基准。数据源自电影等创意视频,其刻意省略某些事件或关系的直接呈现,要求观众进行推断。VRR-QA包含1000个专家标注的问答对,来自1000段创意视频片段,覆盖7个年代、15种类型,涵盖真人与动画作品。对14个主流VideoQA模型的评估显示性能普遍大幅下降,凸显其对表面视觉线索的依赖,并揭示隐性推理的难度。即使表现最好的模型也仅达64%准确率,远低于人类基准。不同模型间性能差异进一步说明任务复杂多样。通过公开数据集与构建框架,VRR-QA为推动视频问答发展提供了严谨、多元且可复现的测试平台。

原文摘要 · Abstract (English)

Video Question Answering (VideoQA) has made significant strides by leveraging multimodal learning to align visual and textual modalities. However, current benchmarks overwhelmingly focus on questions answerable through explicit visual content - actions, objects, and events - directly observable within individual frames or short clips. To truly understand videos as humans do, models must go beyond what is directly shown, inferring hidden relationships and contextual cues that are only implied across frames. Current benchmarks fail to capture this essential aspect of video understanding. To address this gap, we introduce VRR-QA, a benchmark for Visual Relational Reasoning Beyond Explicit Cues. We curate our benchmark from creative and cinematic videos such as movies, that deliberately employ storytelling techniques which omit direct depictions of certain events or relations, requiring viewers to infer them. VRR-QA comprises 1K meticulously expert-annotated QA pairs drawn from 1K creative video clips covering 15 genres across 7 decades of content, from both live-action and animated titles. Our extensive evaluations on 14 leading VideoQA models reveals consistent and significant performance degradation, underscoring their reliance on surface-level visual cues and highlighting the difficulty of implicit reasoning. Even the best model substantially underperforms human baselines with only 64% accuracy. Performance variations across models further illustrate the complexity and diversity of the challenges presented by VRR-QA. By releasing both dataset and data collection framework, VRR-QA establishes a rigorous, diverse, and reproducible testbed for advancing VideoQA: https://swetha5.github.io/ImplicitQA/.

视频理解隐式推理多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。