用全景图连接2D与3D,让视觉语言模型更准定位3D物体。
PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding
- 用带3D语义的全景图作中间表示,适配现有视觉语言模型。
- 在ScanRefer和Nr3D上达到当前最优,对未见数据泛化能力强。
- 适合做3D场景理解、机器人导航的开发者参考。
3D视觉定位(3DVG)是连接视觉语言感知与机器人应用的关键任务,需同时理解语言和进行3D场景推理。传统监督模型依赖显式3D几何信息,但因3D视觉语言数据集稀缺且推理能力弱于现代视觉语言模型(VLMs),泛化性有限。我们提出PanoGrounder框架,将多模态全景图表示与预训练2D VLM结合,实现强视觉语言推理。全景渲染图融合3D语义与几何特征,作为2D与3D间的中间表示,具备两大优势:(i) 可直接输入VLM而无需复杂适配;(ii) 因360度视野保留了长距离物体间关系。设计三阶段流程:根据场景布局选取少量全景视角,用VLM在每张图上定位文本查询,再通过提升融合各视角预测得到单一3D边界框。该方法在ScanRefer和Nr3D上取得领先性能,并在未见3D数据集和文本重述上展现强大泛化能力。
原文摘要 · Abstract (English)
3D Visual Grounding (3DVG) is a critical bridge from vision-language perception to robotics, requiring both language understanding and 3D scene reasoning. Traditional supervised models leverage explicit 3D geometry but exhibit limited generalization, owing to the scarcity of 3D vision-language datasets and the limited reasoning capabilities compared to modern vision-language models (VLMs). We propose a generalizable 3DVG framework, PanoGrounder, that couples multi-modal panoramic representation with pretrained 2D VLMs for strong vision-language reasoning. Panoramic renderings, augmented with 3D semantic and geometric features, serve as an intermediate representation between 2D and 3D, and offer two major benefits: (i) they can be directly fed to VLMs with minimal adaptation and (ii) they retain long-range object-to-object relations thanks to their 360-degree field of view. We devise a three-stage pipeline that places a compact set of panoramic viewpoints considering the scene layout and geometry, grounds a text query on each panoramic rendering with a VLM, and fuses per-view predictions into a single 3D bounding box via lifting. Our approach achieves state-of-the-art results on ScanRefer and Nr3D, and demonstrates strong generalization to unseen 3D datasets and text rephrasings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。