arXiv:2603.25411cs.CV2026-03中稿 · CVPR被引 2

构建分层3D空间理解框架,提升视觉语言模型的立体感知与推理能力。

HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models

  • 分四层递进设计3D空间理解任务,从几何感知到抽象推理
  • 生成500万图像、4500万物体的3D空间VQA数据,用于模型训练
  • 引入点云辅助输入,显著增强模型对空间尺度的理解

实现视觉语言模型(VLMs)的人类级空间智能,需从二维观测中推断三维结构,识别物体在三维空间中的属性与关系,并进行高层次空间推理。本文提出一个系统性的分层框架,将VLMs学习3D空间理解的过程分解为四个逐步复杂化层次:从几何感知到抽象空间推理。基于该框架,我们构建了自动化数据生成流水线,处理约500万张图像及超过4500万物体,生成涵盖多种任务与场景的3D空间问答对,用于VLM的监督微调。同时,我们开发了一种结合RGB-D数据的VLM,以度量级点云图为辅助输入,进一步提升空间理解能力。大量实验表明,本方法在多个空间理解与推理基准上达到当前最优表现,超越专门的空间模型及大型专有系统(如Gemini-2.5-pro和GPT-5)。此外,分析揭示了各层级任务间的明确依赖关系,为多层级任务设计如何促进3D空间智能的涌现提供了新见解。

原文摘要 · Abstract (English)

Achieving human-like spatial intelligence for vision-language models (VLMs) requires inferring 3D structures from 2D observations, recognizing object properties and relations in 3D space, and performing high-level spatial reasoning. In this paper, we propose a principled hierarchical framework that decomposes the learning of 3D spatial understanding in VLMs into four progressively complex levels, from geometric perception to abstract spatial reasoning. Guided by this framework, we construct an automated pipeline that processes approximately 5M images with over 45M objects to generate 3D spatial VQA pairs across diverse tasks and scenes for VLM supervised fine-tuning. We also develop an RGB-D VLM incorporating metric-scale point maps as auxiliary inputs to further enhance spatial understanding. Extensive experiments demonstrate that our approach achieves state-of-the-art performance on multiple spatial understanding and reasoning benchmarks, surpassing specialized spatial models and large proprietary systems such as Gemini-2.5-pro and GPT-5. Moreover, our analysis reveals clear dependencies among hierarchical task levels, offering new insights into how multi-level task design facilitates the emergence of 3D spatial intelligence.

3D理解视觉语言模型空间推理点云融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。