arXiv:2511.05491cs.CV2025-11被引 70

让视觉语言模型学会像人一样感知和推理空间关系。

Visual Spatial Tuning

  • 构建大规模数据集,分阶段训练模型提升空间感知与推理能力。
  • 在多个空间基准上达到34.8%至61.2%的领先性能。
  • 无需牺牲通用能力,适合需要物理理解的AI应用。

捕捉视觉输入中的空间关系是实现类人通用智能的核心。以往研究通过添加专家编码器增强视觉语言模型(VLMs)的空间意识,但带来额外开销且损害通用能力。为在通用架构中提升空间能力,我们提出视觉空间调优(VST),一个从空间感知到推理的完整框架。首先构建包含410万样本、覆盖19种技能的VST-P数据集,涵盖单视图、多图像和视频。随后设计包含13.5万样本的VST-R数据集,引导模型进行空间推理。采用渐进式训练流程:先监督微调建立基础空间知识,再通过强化学习提升空间推理能力。在不损害通用能力的前提下,VST在多个空间基准上持续达到顶尖表现,包括MMSI-Bench的34.8%和VSIBench的61.2%。结果表明,视觉语言动作模型可通过该空间调优范式显著增强,为更贴近物理世界的智能系统铺平道路。

原文摘要 · Abstract (English)

Capturing spatial relationships from visual inputs is a cornerstone of human-like general intelligence. Several previous studies have tried to enhance the spatial awareness of Vision-Language Models (VLMs) by adding extra expert encoders, which brings extra overhead and usually harms general capabilities. To enhance the spatial ability in general architectures, we introduce Visual Spatial Tuning (VST), a comprehensive framework to cultivate VLMs with human-like visuospatial abilities, from spatial perception to reasoning. We first attempt to enhance spatial perception in VLMs by constructing a large-scale dataset termed VST-P, which comprises 4.1 million samples spanning 19 skills across single views, multiple images, and videos. Then, we present VST-R, a curated dataset with 135K samples that instruct models to reason in space. In particular, we adopt a progressive training pipeline: supervised fine-tuning to build foundational spatial knowledge, followed by reinforcement learning to further improve spatial reasoning abilities. Without the side-effect to general capabilities, the proposed VST consistently achieves state-of-the-art results on several spatial benchmarks, including $34.8\%$ on MMSI-Bench and $61.2\%$ on VSIBench. It turns out that the Vision-Language-Action models can be significantly enhanced with the proposed spatial tuning paradigm, paving the way for more physically grounded AI.

视觉语言模型空间推理强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。