arXiv:2603.27437cs.CV2026-03被引 12

通过分层融合视觉、几何与语言信息,提升大模型对3D空间关系的理解能力。

SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning

论文配图:SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
图 1 · 摘自论文原文
  • 分层堆叠多级几何特征,与语言主干同步对齐
  • 在多个3D空间推理基准上达到顶尖性能
  • 适合需要精准空间理解的物理智能系统

大型视觉语言模型(VLMs)在可靠的3D空间推理方面仍表现不足,这是具身与物理人工智能系统的核心挑战。其根源在于难以捕捉精细的3D几何结构与空间关系。尽管近期工作已将多视角几何变换器引入VLM,但通常仅融合视觉与几何编码器的深层特征,忽略了丰富的层级信号,造成空间理解的根本瓶颈。为此,我们提出SpatialStack,一种通用的分层融合框架,可逐步对齐视觉、几何与语言表征在模型层级上的映射。突破传统晚期视觉-几何融合,SpatialStack将多级几何特征与语言主干堆叠并同步,使模型同时具备局部几何精度与全局语义上下文感知能力。基于此框架,我们构建了VLM-SpatialStack,在多个3D空间推理基准上实现领先性能。大量实验与消融分析表明,该多级融合策略持续增强3D理解能力,并在多样化任务中表现出强泛化性,确立SpatialStack作为下一代多模态物理智能系统中视觉-语言-几何融合的有效且可扩展的设计范式。

原文摘要 · Abstract (English)

Large vision-language models (VLMs) still struggle with reliable 3D spatial reasoning, a core capability for embodied and physical AI systems. This limitation arises from their inability to capture fine-grained 3D geometry and spatial relationships. While recent efforts have introduced multi-view geometry transformers into VLMs, they typically fuse only the deep-layer features from vision and geometry encoders, discarding rich hierarchical signals and creating a fundamental bottleneck for spatial understanding. To overcome this, we propose SpatialStack, a general hierarchical fusion framework that progressively aligns vision, geometry, and language representations across the model hierarchy. Moving beyond conventional late-stage vision-geometry fusion, SpatialStack stacks and synchronizes multi-level geometric features with the language backbone, enabling the model to capture both local geometric precision and global contextual semantics. Building upon this framework, we develop VLM-SpatialStack, a model that achieves state-of-the-art performance on multiple 3D spatial reasoning benchmarks. Extensive experiments and ablations demonstrate that our multi-level fusion strategy consistently enhances 3D understanding and generalizes robustly across diverse spatial reasoning tasks, establishing SpatialStack as an effective and extensible design paradigm for vision-language-geometry integration in next-generation multimodal physical AI systems.

3D空间推理多模态融合视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。