arXiv:2509.18905cs.AI2025-09被引 40

评测视觉语言模型的空间智能,发现感知强但推理弱。

How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective

  • 按感知、理解、规划分层构建空间智能评测体系
  • 23个任务设置下,模型在空间规划上表现显著不足
  • 适合研究具身智能与视觉推理的学者参考

视觉空间推理(VSR)是人类核心认知能力,也是推进具身智能和自主系统的关键。尽管视觉语言模型(VLMs)取得进展,但在三维空间表征与推理方面仍远未达到人类水平。本文系统研究了VLMs在空间推理中的表现,涵盖输入模态、模型架构、训练策略与推理机制。将空间智能分为基础感知、空间理解与空间规划三层次,并构建了SIBench基准,整合近20个开源数据集,覆盖23个任务场景。对前沿VLMs的实验显示,模型在基础感知任务中表现良好,但在理解与规划任务中持续落后,尤其在数值估计、多视角推理、时序动态与空间想象方面表现不佳。研究揭示了实现空间智能的重大挑战,同时提供了系统性路线图与全面基准,推动该领域未来发展。相关资源详见https://sibench.github.io/Awesome-Visual-Spatial-Reasoning/

原文摘要 · Abstract (English)

Visual Spatial Reasoning (VSR) is a core human cognitive ability and a critical requirement for advancing embodied intelligence and autonomous systems. Despite recent progress in Vision-Language Models (VLMs), achieving human-level VSR remains highly challenging due to the complexity of representing and reasoning over three-dimensional space. In this paper, we present a systematic investigation of VSR in VLMs, encompassing a review of existing methodologies across input modalities, model architectures, training strategies, and reasoning mechanisms. Furthermore, we categorize spatial intelligence into three levels of capability, ie, basic perception, spatial understanding, spatial planning, and curate SIBench, a spatial intelligence benchmark encompassing nearly 20 open-source datasets across 23 task settings. Experiments with state-of-the-art VLMs reveal a pronounced gap between perception and reasoning, as models show competence in basic perceptual tasks but consistently underperform in understanding and planning tasks, particularly in numerical estimation, multi-view reasoning, temporal dynamics, and spatial imagination. These findings underscore the substantial challenges that remain in achieving spatial intelligence, while providing both a systematic roadmap and a comprehensive benchmark to drive future research in the field. The related resources of this study are accessible at https://sibench.github.io/Awesome-Visual-Spatial-Reasoning/.

空间推理视觉语言模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。