构建视频因果推理新基准,严格测试模型的逐步推断能力。
CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos
- 将视频切分为因果关联片段,强制逐步回答问题
- 含1852个问答对,涵盖6类场景,识别错误类型更精准
- 适合评估视频理解模型的可解释性与推理鲁棒性
大语言模型在文本和图像推理方面取得进展,但视频推理仍面临挑战。现有视频基准多评估浅层理解,允许模型依赖全局上下文,无法严格检验真正的因果与分步推理。我们提出CausalStep,一个面向显式分步因果推理的视频基准。该基准将视频分割为因果关联单元,采用严格的逐步问答协议,要求顺序作答并阻止捷径解法。每个问题均基于错误类型分类设计干扰项,以确保诊断价值。基准包含100个视频、6个类别和1852个多选问答对。我们引入七种诊断指标,实现对因果推理能力的全面评估。对领先商用及开源模型及人类基线的实验显示,当前模型与人类水平的分步推理存在显著差距。CausalStep为推动强健且可解释的视频推理发展提供了严谨评测标准。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have improved reasoning in text and image domains, yet achieving robust video reasoning remains a significant challenge. Existing video benchmarks mainly assess shallow understanding and reasoning and allow models to exploit global context, failing to rigorously evaluate true causal and stepwise reasoning. We present CausalStep, a benchmark designed for explicit stepwise causal reasoning in videos. CausalStep segments videos into causally linked units and enforces a strict stepwise question-answer (QA) protocol, requiring sequential answers and preventing shortcut solutions. Each question includes carefully constructed distractors based on error type taxonomy to ensure diagnostic value. The benchmark features 100 videos across six categories and 1,852 multiple-choice QA pairs. We introduce seven diagnostic metrics for comprehensive evaluation, enabling precise diagnosis of causal reasoning capabilities. Experiments with leading proprietary and open-source models, as well as human baselines, reveal a significant gap between current models and human-level stepwise reasoning. CausalStep provides a rigorous benchmark to drive progress in robust and interpretable video reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。