arXiv:2605.11462cs.CVcs.AI2026-05

用1000万张2D图生成3D空间推理数据,提升视觉语言模型的空间理解能力

SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images

论文配图:SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images
图 1 · 摘自论文原文
  • 将2D图像分解为感知与关系,自动构建深度、布局等结构化监督信号
  • 构建1000万条空间问答对数据集,显著提升模型在多个基准上的空间推理性能
  • 适合希望增强视觉模型空间理解力的研究者或开发者

近期大型视觉语言模型(VLMs)在语义理解上表现卓越,但在空间推理方面仍存在明显短板,难以完成深度排序和坐标定位等基础几何任务。现有方法依赖场景中心的数据集(如多视角扫描或室内视频),但受制于场景数量有限,数据规模与多样性远低于网络级2D图像集合。为此,我们提出SpatialForge,一种可扩展的数据合成管道,将真实世界2D图像转化为空间推理监督信号。该方法将空间推理拆解为感知与关系两部分,构建覆盖深度、布局及视角依赖推理的结构化标注,并通过自动验证保证数据质量。基于此,我们构建了包含1000万条空间问答对的SpatialForge-10M数据集。大量实验表明,在该数据集上训练的标准VLMs在多个空间推理基准上表现显著提升,证明利用大规模2D数据强化3D感知空间推理的有效性。

原文摘要 · Abstract (English)

Recent advancements in Large Vision-Language Models (VLMs) have demonstrated exceptional semantic understanding, yet these models consistently struggle with spatial reasoning, often failing at fundamental geometric tasks such as depth ordering and precise coordinate grounding. Recent efforts introduce spatial supervision from scene-centric datasets (e.g., multi-view scans or indoor video), but are constrained by the limited number of underlying scenes. As a result, the scale and diversity of such data remain significantly smaller than those of web-scale 2D image collections. To address this limitation, we propose SpatialForge, a scalable data synthesis pipeline that transforms in-the-wild 2D images into spatial reasoning supervision. Our approach decomposes spatial reasoning into perception and relation, and constructs structured supervision signals covering depth, layout, and viewpoint-dependent reasoning, with automatic verification to ensure data quality. Based on this pipeline, we build SpatialForge-10M, a large-scale dataset containing 10 million spatial QA pairs. Extensive experiments across multiple spatial reasoning benchmarks demonstrate that training on SpatialForge-10M significantly improves the spatial reasoning ability of standard VLMs, highlighting the effectiveness of scaling 2D data for 3D-aware spatial reasoning.

空间推理视觉语言模型数据合成3D理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。