首个全面评估视频空间智能的基准,揭示大模型与人类差距超60%。
MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence
- 构建四层视频空间理解框架,覆盖感知、规划、预测与跨视频推理。
- 涵盖1278段视频、1106个问题,由3D视觉专家人工标注并提供推理依据。
- 适合研究多模态大模型空间推理能力的学者,尤其关注机器人与场景理解。
连续视觉输入中的空间理解对多模态大模型成为物理环境通用助手至关重要,但目前尚无全面评估该能力的基准。本文提出MMSI-Video-Bench,一个完全人工标注的视频空间智能基准。它基于四层框架(感知、规划、预测、跨视频推理),包含来自25个数据集及自建视频的1,278段视频和1,106个问题,每道题均由3D视觉专家设计并提供解释性理由,确保精准定位。该基准支持三个领域子基准(室内场景感知、机器人、定位),可针对性评估模型能力。我们评估了25个主流开源与专有模型,发现多数模型表现接近随机水平,最佳模型仍落后人类近60%。空间微调模型在本基准上仍无法有效泛化。细粒度错误分析揭示几何推理、运动定位、长时序预测与跨视频对应等系统性缺陷。典型帧采样策略在此推理密集型任务中表现不佳,且3D空间线索与思维链提示均未带来显著提升。本基准有望推动视频空间智能的发展。
原文摘要 · Abstract (English)
Spatial understanding over continuous visual input is crucial for MLLMs to evolve into general-purpose assistants in physical environments. Yet there is still no comprehensive benchmark that holistically assesses the progress toward this goal. In this work, we introduce MMSI-Video-Bench, a fully human-annotated benchmark for video-based spatial intelligence in MLLMs. It operationalizes a four-level framework, Perception, Planning, Prediction, and Cross-Video Reasoning, through 1,106 questions grounded in 1,278 clips from 25 datasets and in-house videos. Each item is carefully designed and reviewed by 3DV experts with explanatory rationales to ensure precise, unambiguous grounding. Leveraging its diverse data sources and holistic task coverage, MMSI-Video-Bench also supports three domain-oriented sub-benchmarks (Indoor Scene Perception Bench, Robot Bench and Grounding Bench) for targeted capability assessment. We evaluate 25 strong open-source and proprietary MLLMs, revealing a striking human--AI gap: many models perform near chance, and the best reasoning model lags humans by nearly 60%. We further find that spatially fine-tuned models still fail to generalize effectively on our benchmark. Fine-grained error analysis exposes systematic failures in geometric reasoning, motion grounding, long-horizon prediction, and cross-video correspondence. We also show that typical frame-sampling strategies transfer poorly to our reasoning-intensive benchmark, and that neither 3D spatial cues nor chain-of-thought prompting yields meaningful gains. We expect our benchmark to establish a solid testbed for advancing video-based spatial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。