提出视频空间因果预测新任务,评估模型推断未见时空状态能力
SCP: Spatial Causal Prediction in Video
- 设计新任务SCP,要求模型推断视频中未观测的过去或未来空间状态
- 构建含2500个问答对的SCP-Bench基准,覆盖1181段多视角视频
- 发现当前模型在时空外推和因果推理上远逊人类,适合研究视觉推理者关注
空间推理——理解空间关系、因果性及动态演化——是人类智能的核心,也是自动驾驶与机器人等现实应用的关键。现有研究主要评估模型对可见时空信息的理解,忽视其推断未见时空状态的能力。本文提出空间因果预测(SCP)新任务,挑战模型超越观察范围,预测空间因果结果。我们构建了包含2,500个问答对的SCP-Bench基准,涵盖1,181段视频,覆盖多样视角、场景与因果方向,支持系统化评估。在23个先进模型上进行综合实验,揭示人类与模型间存在显著性能差距,模型在时间外推和因果建模方面表现薄弱。进一步分析影响性能的关键因素,提出感知增强与推理引导策略,以推动空间因果智能发展。
原文摘要 · Abstract (English)
Spatial reasoning, the ability to understand spatial relations, causality, and dynamic evolution, is central to human intelligence and essential for real-world applications such as autonomous driving and robotics. Existing studies, however, primarily assess models on visible spatio-temporal understanding, overlooking their ability to infer unseen past or future spatial states. In this work, we introduce Spatial Causal Prediction (SCP), a new task paradigm that challenges models to reason beyond observation and predict spatial causal outcomes. We further construct SCP-Bench, a benchmark comprising 2,500 QA pairs across 1,181 videos spanning diverse viewpoints, scenes, and causal directions, to support systematic evaluation. Through comprehensive experiments on {23} state-of-the-art models, we reveal substantial gaps between human and model performance, limited temporal extrapolation, and weak causal grounding. We further analyze key factors influencing performance and propose perception-enhancement and reasoning-guided strategies toward advancing spatial causal intelligence. The project page is https://guangstrip.github.io/SCP-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。