arXiv:2602.03916cs.CVcs.CE2026-02中稿 · ICLR被引 10

构建真实场景下的空间推理评测基准,揭示视觉语言模型的显著能力短板。

SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?

  • 设计6大类30种真实场景任务,覆盖相对位置、深度遮挡等复杂空间关系
  • 多模型测试显示最高准确率仅54.93%,远低于人类87.57%的水平
  • 适合关注视觉语言模型空间理解能力提升的研究者使用

空间推理是人类认知的核心能力,但对当前视觉语言模型(VLMs)仍是重大挑战。以往研究依赖合成或大模型生成环境,任务设计单一且呈谜题化,难以反映真实世界中的复杂性、视觉噪声和多样空间关系。为此,我们提出SpatiaLab,一个面向真实、无约束情境下VLM空间推理能力的综合性评测基准。SpatiaLab包含1,400个视觉问答对,涵盖相对位置、深度与遮挡、方向、大小与尺度、空间导航和3D几何六大类别,每类含五个子类,共30种任务类型。每个子类至少25个问题,每大类不少于200个问题,支持多选与开放回答两种评估形式。在多种先进VLM上进行实验,包括开源与闭源模型、专注推理及专用空间推理模型,结果显示其空间推理能力与人类存在显著差距:多选设置下,InternVL3.5-72B最高准确率为54.93%,人类为87.57%;开放回答中,所有模型性能下降约10-25%,GPT-5-mini最高达40.93%,人类为64.93%。结果表明,模型在处理复杂空间关系、深度感知、导航与3D几何方面仍存关键缺陷。SpatiaLab提供多样化真实世界评估框架,揭示了推动VLM实现稳健、人类对齐空间理解的关键挑战与机遇,为未来研究提供重要基准。项目地址:https://spatialab-reasoning.github.io/

原文摘要 · Abstract (English)

Spatial reasoning is a fundamental aspect of human cognition, yet it remains a major challenge for contemporary vision-language models (VLMs). Prior work largely relied on synthetic or LLM-generated environments with limited task designs and puzzle-like setups, failing to capture the real-world complexity, visual noise, and diverse spatial relationships that VLMs encounter. To address this, we introduce SpatiaLab, a comprehensive benchmark for evaluating VLMs' spatial reasoning in realistic, unconstrained contexts. SpatiaLab comprises 1,400 visual question-answer pairs across six major categories: Relative Positioning, Depth & Occlusion, Orientation, Size & Scale, Spatial Navigation, and 3D Geometry, each with five subcategories, yielding 30 distinct task types. Each subcategory contains at least 25 questions, and each main category includes at least 200 questions, supporting both multiple-choice and open-ended evaluation. Experiments across diverse state-of-the-art VLMs, including open- and closed-source models, reasoning-focused, and specialized spatial reasoning models, reveal a substantial gap in spatial reasoning capabilities compared with humans. In the multiple-choice setup, InternVL3.5-72B achieves 54.93% accuracy versus 87.57% for humans. In the open-ended setting, all models show a performance drop of around 10-25%, with GPT-5-mini scoring highest at 40.93% versus 64.93% for humans. These results highlight key limitations in handling complex spatial relationships, depth perception, navigation, and 3D geometry. By providing a diverse, real-world evaluation framework, SpatiaLab exposes critical challenges and opportunities for advancing VLMs' spatial reasoning, offering a benchmark to guide future research toward robust, human-aligned spatial understanding. SpatiaLab is available at: https://spatialab-reasoning.github.io/.

空间推理视觉语言模型评测基准3D理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。