arXiv:2512.16561cs.CV2025-12被引 21

让视觉语言模型直接理解3D空间关系,提升对物体位置的精准定位能力。

N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models

  • 引入原生3D感知机制,直接根据文字描述定位物体在3D空间中的位置。
  • 构建大规模3D标注数据集,规模超过现有最大单图3D检测数据集六倍以上。
  • 支持链式思维推理,适合需要深度空间理解的机器人、自动驾驶等应用。

当前多模态模型虽能基于2D图像回答问题,但缺乏内在的3D物体感知能力,限制了其对3D场景中空间关系和深度线索的理解。本文提出N3D-VLM,一种统一框架,将原生3D物体感知与3D感知推理无缝结合,实现精确的3D定位与可解释的空间理解。不同于直接从RGB/RGB-D输入端到端预测答案的传统方法,本方法赋予模型原生3D感知能力,使其能根据文本描述直接定位物体在3D空间中的位置。基于精准的3D定位,模型进一步执行显式的3D推理,获得更可解释、结构化的空间理解。为支持该能力的稳健训练,我们设计了一个可扩展的数据构建流程,利用深度估计将大规模2D标注迁移到3D空间,显著提升3D物体定位数据的多样性和覆盖范围,生成的数据量超过现有最大单图3D检测数据集六倍以上。此外,该流程还生成面向链式思维(CoT)推理的3D空间问答数据集,支持3D物体定位与3D空间推理的联合训练。实验表明,该统一框架不仅在3D定位任务上达到顶尖性能,且在视觉语言模型的3D空间推理能力上持续超越现有方法。

原文摘要 · Abstract (English)

While current multimodal models can answer questions based on 2D images, they lack intrinsic 3D object perception, limiting their ability to comprehend spatial relationships and depth cues in 3D scenes. In this work, we propose N3D-VLM, a novel unified framework that seamlessly integrates native 3D object perception with 3D-aware visual reasoning, enabling both precise 3D grounding and interpretable spatial understanding. Unlike conventional end-to-end models that directly predict answers from RGB/RGB-D inputs, our approach equips the model with native 3D object perception capabilities, enabling it to directly localize objects in 3D space based on textual descriptions. Building upon accurate 3D object localization, the model further performs explicit reasoning in 3D, achieving more interpretable and structured spatial understanding. To support robust training for these capabilities, we develop a scalable data construction pipeline that leverages depth estimation to lift large-scale 2D annotations into 3D space, significantly increasing the diversity and coverage for 3D object grounding data, yielding over six times larger than the largest existing single-image 3D detection dataset. Moreover, the pipeline generates spatial question-answering datasets that target chain-of-thought (CoT) reasoning in 3D, facilitating joint training for both 3D object localization and 3D spatial reasoning. Experimental results demonstrate that our unified framework not only achieves state-of-the-art performance on 3D grounding tasks, but also consistently surpasses existing methods in 3D spatial reasoning in vision-language model.

3D感知空间推理视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。