用多视角图像自监督学习3D空间结构,让视觉语言模型更懂三维场景。
Spa3R: Predictive Spatial Field Modeling for 3D Visual Reasoning

- 从无相机参数的多视角图中学习不变的空间表征
- 在新视角上预测几何与语义特征场,平均准确率58.6%
- 适合需要三维推理能力的视觉语言模型研究者
视觉语言模型在2D视觉理解上表现优异,但在3D空间推理方面仍受限。现有方法或依赖显式3D模态(影响扩展性),或注入部分视图相关的几何先验,让语言模型从稀疏线索中恢复全局场景结构。本文提出Spa3R,一种自监督框架,从无相机参数的多视角RGB图像中学习统一、视图无关的空间表征。其预测性空间场建模目标将上下文视图压缩为紧凑潜在表示,并在新视角上预测对齐的几何与语义特征场,从而促进场景几何与布局的一致编码。我们将预训练的Spa3R编码器通过轻量级残差交叉注意力适配器集成至视觉语言模型,得到Spa3-VLM,使语言推理基于全局空间上下文。Spa3-VLM在VSI-Bench上取得58.6%的平均得分,在三个额外的空间推理基准上达到领先或竞争性能。结果表明,预测性空间表征学习为3D推理提供了有效的视觉基础。
原文摘要 · Abstract (English)
Vision-language models excel at 2D visual understanding but remain limited in 3D spatial reasoning. Existing approaches either depend on explicit 3D modalities, which limits scalability, or inject partial, view-conditioned geometric priors and leave the language model to recover global scene structure from sparse cues. We introduce Spa3R, a self-supervised framework that learns a unified, view-invariant spatial representation from unposed multi-view RGB images. Its Predictive Spatial Field Modeling objective compresses context views into a compact latent representation and predicts aligned geometric and semantic feature fields at novel viewpoints, thereby encouraging coherent encoding of scene geometry and layout. We integrate the pre-trained Spa3R Encoder into a vision-language model through a lightweight residual cross-attention adapter, yielding Spa3-VLM and grounding language reasoning in global spatial context. Spa3-VLM achieves an average score of 58.6% on VSI-Bench and delivers leading or competitive performance across three additional spatial reasoning benchmarks. These results demonstrate that predictive spatial representation learning provides an effective visual foundation for 3D reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。