arXiv:2503.22976cs.CV2025-03NeurIPS被引 117

用2D图像数据训练视觉语言模型,提升其3D空间感知与推理能力。

From Flatland to Space: Teaching Vision-Language Models to Perceive and Reason in 3D

  • 基于带3D真值的场景数据生成多样化的2D空间任务数据
  • 构建700万样本的SPAR-7M数据集,在2D空间基准上达顶尖性能
  • 提出支持单/多视角的SPAR-Bench评测基准,适合研究3D理解的学者

近年来,视觉语言模型在视觉-语言理解方面取得进展,但仍难以处理空间感知问题,限制了其对复杂3D场景的推理能力。不同于将3D表示引入模型的方法,本文通过利用具有空间相关性的2D图像数据来挖掘视觉语言模型的潜力。为此,我们设计了一套基于带有3D真值场景数据的2D空间数据生成与标注流程,可创建从基础感知到复杂推理的多样化空间任务。基于此流程,我们构建了大规模数据集SPAR-7M,涵盖数千个来自多个公开数据集的场景。同时,我们提出了SPAR-Bench,一个比现有空间评测更全面的基准,支持单视图与多视图输入。在SPAR-7M与大规模2D数据集上联合训练后,模型在2D空间基准上达到当前最优表现;进一步在3D特定任务数据集上微调,仍取得有竞争力的结果,验证了该数据集在提升空间推理能力方面的有效性。

原文摘要 · Abstract (English)

Recent advances in LVLMs have improved vision-language understanding, but they still struggle with spatial perception, limiting their ability to reason about complex 3D scenes. Unlike previous approaches that incorporate 3D representations into models to improve spatial understanding, we aim to unlock the potential of VLMs by leveraging spatially relevant image data. To this end, we introduce a novel 2D spatial data generation and annotation pipeline built upon scene data with 3D ground-truth. This pipeline enables the creation of a diverse set of spatial tasks, ranging from basic perception tasks to more complex reasoning tasks. Leveraging this pipeline, we construct SPAR-7M, a large-scale dataset generated from thousands of scenes across multiple public datasets. In addition, we introduce SPAR-Bench, a benchmark designed to offer a more comprehensive evaluation of spatial capabilities compared to existing spatial benchmarks, supporting both single-view and multi-view inputs. Training on both SPAR-7M and large-scale 2D datasets enables our models to achieve state-of-the-art performance on 2D spatial benchmarks. Further fine-tuning on 3D task-specific datasets yields competitive results, underscoring the effectiveness of our dataset in enhancing spatial reasoning.

3D感知视觉语言模型空间推理数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。