arXiv:2506.09943cs.CVcs.AI2025-06被引 32

构建真实物理场景下的因果推理视频问答基准,检验模型预测行为后果能力。

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models

  • 设计五类因果问题:反事实、假设、预见、规划与描述,覆盖真实场景行为推演。
  • 当前顶尖多模态模型在预见与假设类问题上表现远低于人类,差距显著。
  • 通过质量控制防止模型依赖语言捷径,强制其理解深层视觉与物理规律。

我们提出 CausalVQA,一个用于视频问答(VQA)的因果推理基准数据集,包含五类问题:反事实、假设、预见、规划与描述,旨在考察模型对真实世界物理因果关系的理解。现有 VQA 基准或侧重表面感知,或局限于仿真环境中的狭义物理推理。CausalVQA 填补了这一空白,聚焦真实场景中对不同行为与事件可能结果的预测能力。我们设计了质量控制机制,防止模型利用语言捷径,要求其基于深度视觉理解作答。实验表明,当前前沿多模态模型在该基准上表现明显落后于人类,尤其在预见和假设类问题上差距巨大,凸显其在时空推理、物理原理理解和潜在替代方案认知方面的不足。

原文摘要 · Abstract (English)

We introduce CausalVQA, a benchmark dataset for video question answering (VQA) composed of question-answer pairs that probe models' understanding of causality in the physical world. Existing VQA benchmarks either tend to focus on surface perceptual understanding of real-world videos, or on narrow physical reasoning questions created using simulation environments. CausalVQA fills an important gap by presenting challenging questions that are grounded in real-world scenarios, while focusing on models' ability to predict the likely outcomes of different actions and events through five question types: counterfactual, hypothetical, anticipation, planning and descriptive. We designed quality control mechanisms that prevent models from exploiting trivial shortcuts, requiring models to base their answers on deep visual understanding instead of linguistic cues. We find that current frontier multimodal models fall substantially below human performance on the benchmark, especially on anticipation and hypothetical questions. This highlights a challenge for current systems to leverage spatial-temporal reasoning, understanding of physical principles, and comprehension of possible alternatives to make accurate predictions in real-world settings.

因果推理视频问答物理理解多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。