arXiv:2412.04383cs.CVcs.RO2024-12CVPR被引 82

用2D视觉语言模型实现零样本3D物体定位,无需标注数据

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

  • 将3D场景转为2D图像与空间文本混合表示,适配2D模型输入
  • 在ScanRefer和Nr3D上分别领先前序方法7.7%和7.1%
  • 无需标注数据即可定位未见类别,适合新场景快速部署

3D视觉定位(3DVG)旨在根据文本描述定位3D场景中的物体,对增强现实和机器人应用至关重要。传统方法依赖标注的3D数据集和预定义类别,限制了可扩展性与适应性。为此,我们提出SeeGround,一种利用大规模2D数据训练的2D视觉语言模型(VLMs)的零样本3DVG框架。SeeGround将3D场景表示为查询对齐的渲染图像与空间增强文本描述的混合形式,弥合3D数据与2D-VLM输入格式之间的差距。提出两个模块:视角适配模块动态选择与查询相关的视角进行图像渲染;融合对齐模块将2D图像与3D空间描述融合,提升定位精度。在ScanRefer和Nr3D上的大量实验表明,该方法显著优于现有零样本方法,显著超越弱监督方法,并媲美部分全监督方法,在ScanRefer上领先7.7%,在Nr3D上领先7.1%,展现了在复杂3DVG任务中的有效性。

原文摘要 · Abstract (English)

3D Visual Grounding (3DVG) aims to locate objects in 3D scenes based on textual descriptions, essential for applications like augmented reality and robotics. Traditional 3DVG approaches rely on annotated 3D datasets and predefined object categories, limiting scalability and adaptability. To overcome these limitations, we introduce SeeGround, a zero-shot 3DVG framework leveraging 2D Vision-Language Models (VLMs) trained on large-scale 2D data. SeeGround represents 3D scenes as a hybrid of query-aligned rendered images and spatially enriched text descriptions, bridging the gap between 3D data and 2D-VLMs input formats. We propose two modules: the Perspective Adaptation Module, which dynamically selects viewpoints for query-relevant image rendering, and the Fusion Alignment Module, which integrates 2D images with 3D spatial descriptions to enhance object localization. Extensive experiments on ScanRefer and Nr3D demonstrate that our approach outperforms existing zero-shot methods by large margins. Notably, we exceed weakly supervised methods and rival some fully supervised ones, outperforming previous SOTA by 7.7% on ScanRefer and 7.1% on Nr3D, showcasing its effectiveness in complex 3DVG tasks.

3D视觉定位零样本学习视觉语言模型开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。