构建视频因果推理新基准,要求模型定位多阶段时空证据链。
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering

- 通过人机协作生成2066个因果问题,标注时间片段与轨迹框
- 当前视觉语言模型在因果推理上准确率低,难构建精确证据链
- 适合研究视频理解、可解释性及因果推理的学者使用
视频中的因果推理是视觉语言模型(VLMs)的重要挑战,需超越表层感知,深入理解因果机制。现有基准缺乏细粒度且可落地的证据支持,难以严格评估该能力。为此,我们提出CaST-Bench,一个基于因果链的时空视频推理基准。该基准包含1,015个视频上的2,066个复杂因果问题,要求模型识别并定位多个时空证据组成的因果链条。通过人机协作流程,所有因果链均以时间片段和边界框轨迹标注。我们还设计了综合评估体系,引入新指标,不仅评估答案正确性,更检验视觉证据的可追溯性。这种可追溯性有助于减少虚假关联、提升模型可信度。实验表明,当前VLMs在因果问题上表现不佳,主要受限于构建精确、可落地因果链的能力。这揭示了未来VLMs改进的关键方向。
原文摘要 · Abstract (English)
Cause-and-effect reasoning in video is a significant challenge for Vision-Language Models (VLMs), as it requires going beyond surface-level perception to a deeper understanding of causal mechanisms. However, existing benchmarks rarely provide the fine-grained, grounded evidence needed to rigorously evaluate this capability. To address this gap, we introduce CaST-Bench, a benchmark for Causal Chain-Grounded Spatio-Temporal Video Reasoning. CaST-Bench presents complex causal questions that require models to identify and localize a chain of multiple spatio-temporal evidences. Through a human-AI collaborative pipeline, we construct a high-quality dataset of 2,066 questions over 1,015 videos, with causal chains annotated by temporal segments and bounding-box tracks. Furthermore, we design a comprehensive evaluation suite with novel metrics that assess not only answer correctness but also the capability for visual evidence grounded reasoning. This grounding is crucial for improving accuracy by mitigating spurious correlations and for enhancing user trust by making models more transparent. Our experiments show that current VLMs struggle with causal questions, largely due to their limited ability to construct precise and grounded causal chains. This highlights an important direction for improving future VLMs. Homepage: https://woven-by-toyota.github.io/CaST-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。