测试大模型能否像人一样从视频中理解空间关系
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

- 构建5000+题的视觉空间智能评测集
- 模型能部分生成认知地图,但空间推理仍弱于人类
- 显式画地图比语言推理更有效提升空间能力
人类具备从序列视觉观察中记忆空间的能力。然而,基于百万级视频数据训练的多模态大模型(MLLMs)是否也能‘在空间中思考’?我们提出一个包含超过5000个问答对的新型视频型视觉-空间智能基准(VSI-Bench),发现MLLMs展现出具有竞争力但低于人类水平的视觉-空间智能。我们探究模型在语言和视觉层面如何表达空间思维,发现尽管空间推理仍是模型性能提升的主要瓶颈,局部世界模型和空间意识已在模型中出现。值得注意的是,主流语言推理方法(如思维链、自一致性、思维树)未能提升性能,而显式生成认知地图则显著增强了模型的空间距离判断能力。
原文摘要 · Abstract (English)
Humans possess the visual-spatial intelligence to remember spaces from sequential visual observations. However, can Multimodal Large Language Models (MLLMs) trained on million-scale video datasets also ``think in space'' from videos? We present a novel video-based visual-spatial intelligence benchmark (VSI-Bench) of over 5,000 question-answer pairs, and find that MLLMs exhibit competitive - though subhuman - visual-spatial intelligence. We probe models to express how they think in space both linguistically and visually and find that while spatial reasoning capabilities remain the primary bottleneck for MLLMs to reach higher benchmark performance, local world models and spatial awareness do emerge within these models. Notably, prevailing linguistic reasoning techniques (e.g., chain-of-thought, self-consistency, tree-of-thoughts) fail to improve performance, whereas explicitly generating cognitive maps during question-answering enhances MLLMs' spatial distance ability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。