arXiv:2509.17664cs.CVcs.AI2025-09NeurIPS被引 17

让视觉语言模型学会精准量空间关系,提升3D理解能力。

SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models

  • 用深度位置编码增强模型对空间的感知能力
  • 构建含70万问答对的大型空间测量数据集
  • 在多个基准上超越GPT-4o和Intern-VL3-78B

尽管视觉语言模型(VLMs)在二维语义理解方面表现优异,但其对三维空间关系的定量推理能力仍较弱,主要受限于2D图像的空间表征能力不足。本文分析了制约VLM空间理解的关键问题,提出SD-VLM框架,通过两项关键贡献显著提升VLM的基础空间感知能力:(1)构建大规模空间测量与理解(MSMU)数据集,包含70万组问答对、250万条物理数值标注及1万条链式思维增强样本;(2)引入一种简洁的深度位置编码方法,强化VLM的空间意识。我们训练了强大的通用型模型SD-VLM,其在自建的MSMU-Bench上达到领先水平,并在Q-Spatial与SpatialRGPT-Bench等其他空间理解基准上展现出良好泛化能力。大量实验表明,SD-VLM在MSMU-Bench上优于GPT-4o和Intern-VL3-78B分别达26.91%和25.56%。代码与模型已开源。

原文摘要 · Abstract (English)

While vision language models (VLMs) excel in 2D semantic visual understanding, their ability to quantitatively reason about 3D spatial relationships remains under-explored, due to the deficiency of 2D images' spatial representation ability. In this paper, we analyze the problem hindering VLMs' spatial understanding abilities and propose SD-VLM, a novel framework that significantly enhances fundamental spatial perception abilities of VLMs through two key contributions: (1) propose Massive Spatial Measuring and Understanding (MSMU) dataset with precise spatial annotations, and (2) introduce a simple depth positional encoding method strengthening VLMs' spatial awareness. MSMU dataset covers massive quantitative spatial tasks with 700K QA pairs, 2.5M physical numerical annotations, and 10K chain-of-thought augmented samples. We have trained SD-VLM, a strong generalist VLM which shows superior quantitative spatial measuring and understanding capability. SD-VLM not only achieves state-of-the-art performance on our proposed MSMU-Bench, but also shows spatial generalization abilities on other spatial understanding benchmarks including Q-Spatial and SpatialRGPT-Bench. Extensive experiments demonstrate that SD-VLM outperforms GPT-4o and Intern-VL3-78B by 26.91% and 25.56% respectively on MSMU-Bench. Code and models are released at https://github.com/cpystan/SD-VLM.

空间理解视觉语言模型深度编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。