arXiv:2607.20868cs.CV2026-07被引 1

评测大模型能否从动态画面中进行直观推理,发现现有模型仍远不如人类。

ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?

论文配图:ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?
图 1 · 摘自论文原文
  • 构建新基准ViSTR-Bench,用连续视频评估模型的时空推理能力。
  • 15个子任务、1340组问答数据,覆盖运动感知、空间关系等四维度。
  • 多数大模型在复杂动态推理上表现不佳,距离人类水平仍有差距。

多模态大语言模型(MLLMs)在众多专家级任务中取得显著进展,但在人类通过持续观察现实世界自然掌握的基础能力上仍显不足,如空间感知和动态推理。尽管已有研究意识到这一差距并推出专门基准,但现有评测多聚焦静态场景或需精确数值预测,对基于时间线索的直觉推理关注不足。本文提出视觉时空推理基准(ViSTR-Bench),旨在系统评估MLLMs是否能从动态场景的连续视觉线索中进行定性推理。该基准遵循时间强调、推理导向与定性评价原则,涵盖运动感知、空间关系、结果预测与物理动力学四个维度,包含15个子任务及1,340组高质量视频问答对,覆盖桌面、室内与室外多种场景。对主流商业、开源及专用空间类MLLM的广泛评估表明,尽管具备较强的通用视频理解能力,当前模型在复杂时空推理上仍存在明显瓶颈,远未达到人类水平。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental abilities that humans naturally develop through continuous observation of the real world, such as spatial perception and dynamic reasoning. Recent studies have recognized this gap and introduced dedicated benchmarks to evaluate the spatial-temporal capabilities of MLLMs. However, existing benchmarks mostly focus on static scenes or require exact quantitative predictions, leaving intuitive reasoning from temporal cues largely underexplored. In this paper, we introduce the Visual Spatial-Temporal Reasoning Benchmark (ViSTR-Bench), a novel evaluation suite designed to systematically assess whether MLLMs can perform qualitative reasoning from continuous visual cues in dynamic scenes. Guided by the principles of temporal emphasis, reasoning orientation, and qualitative evaluation, ViSTR-Bench establishes a comprehensive four-dimensional evaluations covering Motion Perception, Spatial Relations, Outcome Prediction, and Physical Dynamics. The benchmark comprises 15 distinct subtasks and 1,340 high-quality video question-answer pairs spanning diverse tabletop, indoor, and outdoor scenarios. Extensive evaluations of a broad spectrum of state-of-the-art proprietary, open-source, and specialized spatial MLLMs reveal that, despite their strong general video understanding capabilities, current models still face substantial bottlenecks in complex spatial-temporal reasoning and remain far below human performance.

多模态动态推理视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。