arXiv:2511.13782cs.AI2025-11被引 1

提出新基准与框架,揭示视觉语言模型空间推理的瓶颈与改进路径

Imagine in Space: Exploring the Frontier of Spatial Intelligence and Reasoning Efficiency in Vision Language Models

  • 用合成数据构建空间推理评测基准,量化模型表现与效率
  • 发现主流模型依赖语言而非视觉特征,且生成耗时随复杂度飙升
  • 提出图像驱动框架,隐式构建内部世界模型提升空间推理能力

大型语言模型(LLMs)和视觉语言模型(VLMs),如DeepSeek R1、OpenAI o3和Gemini 2.5 Pro,已在逻辑推理、问题求解和决策中展现强大能力。然而,空间推理——人类认知的基础能力,包括心理旋转、导航和空间关系理解——仍是当前先进VLM的重大挑战。我们假设:想象,即对空间状态的内在模拟,是空间世界模型中的主导推理机制。为验证该假设并系统探究现有VLM的空间推理机制,我们引入SpatiaLite,一个完全合成的基准,联合衡量空间推理准确率与推理效率。综合实验揭示三个关键发现:第一,先进VLM主要依赖语言表示进行推理与想象,导致在需要感知空间关系和三维几何变换(如心理旋转或投影预测)的视觉中心任务中表现严重不足;第二,当前空间推理机制存在严重低效,令牌使用量随变换复杂度急剧上升;第三,我们提出图像驱动框架(IDF)用于数据合成与训练,可隐式构建对空间推理至关重要的内部世界模型。基于SpatiaLite,本工作刻画了先进VLM的空间推理边界与模式,识别关键短板,并为未来突破提供方向。

原文摘要 · Abstract (English)

Large language models (LLMs) and vision language models (VLMs), such as DeepSeek R1,OpenAI o3, and Gemini 2.5 Pro, have demonstrated remarkable reasoning capabilities across logical inference, problem solving, and decision making. However, spatial reasoning:a fundamental component of human cognition that includes mental rotation, navigation, and spatial relationship comprehension remains a significant challenge for current advanced VLMs. We hypothesize that imagination, the internal simulation of spatial states, is the dominant reasoning mechanism within a spatial world model. To test this hypothesis and systematically probe current VLM spatial reasoning mechanisms, we introduce SpatiaLite, a fully synthetic benchmark that jointly measures spatial reasoning accuracy and reasoning efficiency. Comprehensive experiments reveal three key findings. First, advanced VLMs predominantly rely on linguistic representations for reasoning and imagination, resulting in significant deficiencies on visual centric tasks that demand perceptual spatial relations and 3D geometry transformations such as mental rotation or projection prediction. Second, advanced VLMs exhibit severe inefficiency in their current spatial reasoning mechanisms, with token usage growing rapidly as transformation complexity increases. Third, we propose an Imagery Driven Framework (IDF) for data synthesis and training, which can implicitly construct an internal world model that is critical for spatial reasoning in VLMs. Building on SpatiaLite, this work delineates the spatial reasoning limits and patterns of advanced VLMs, identifies key shortcomings, and informs future advances

空间推理视觉语言模型世界模型高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。