arXiv:2608.01899cs.CVcs.CL2026-08

让视觉语言模型学会物理空间推理,不加额外模块也能显著提升表现。

SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models

论文配图:SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models
图 1 · 摘自论文原文
  • 用伪深度和相机信息做监督,激活模型内在空间知识。
  • 在VSI-Bench上达71.6分,首个突破70分的模型。
  • 轻量插件式设计,不破坏原模型通用能力,适合多场景应用。

视觉语言模型在常识推理任务中表现良好,但在视觉空间推理上存在不足。现有方法通常引入额外的3D先验输入或外部空间编码器,增加了复杂性,并在空间微调后削弱了模型的通用能力。为此,我们提出一种参数高效的时空视觉语言模型(SpatioLM),在无需额外3D先验或第三方空间编码器的情况下增强空间智能。具体而言,设计了一个即插即用、非侵入式的时空视觉模块,激发模型内在的空间知识;同时创新性地利用伪深度与相机信息作为监督信号,引导模型学习物理一致的表示。大量实验表明,SpatioLM在多种任务中实现显著提升,包括空间感知与理解,且有效抑制通用能力下降。特别地,模型在VSI-Bench上取得71.6分,为首个突破70分的模型。此外,在具身操作任务迁移中也表现出色。代码已开源。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduce extra 3D prior inputs or external spatial encoders, which increase complexity and degrade the underlying VLMs' general-purpose capabilities after spatial fine-tuning. To this end, we propose a parameter-efficient \textit{\textbf{Spatio}-vision \textbf{L}anguage \textbf{M}odels (SpatioLM)}, that enhances spatial intelligence without extra 3D prior inputs or third-party spatial encoders. Concretely, we design a plug-and-play and non-invasive spatio-vision module that elicits the spatial knowledge inherent in VLMs. Furthermore, we innovatively leverage pseudo depth and camera information as supervision to guide the model in learning physically coherent representations. Extensive experiments show that SpatioLM achieves significant improvements in diverse tasks, including spatial perception and understanding while effectively limiting the degradation of general capabilities. Notably, the model achieves an impressive score of 71.6 on the VSI-Bench (the first model to surpass 70). In addition, it attains competitive performance when transferred to embodied manipulation tasks. Code is available at \href{https://github.com/xiaomi-research/spatio-lm}{\faGithub~spatio-lm}.

视觉语言模型空间推理物理智能轻量设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。