用2D视觉语言模型实现零样本3D物体定位,无需3D训练数据
Zero-Shot 3D Visual Grounding from Vision-Language Models
- 通过渲染视图与空间文本结合,打通2D与3D模态差距
- 在ScanRefer和Nr3D上分别领先基线7.7%和7.1%,接近有监督模型表现
- 适合做开放世界机器人、AR等需泛化定位的场景
3D视觉定位(3DVG)旨在利用自然语言描述在3D场景中定位目标物体,支持增强现实、机器人等下游应用。现有方法通常依赖标注的3D数据和预定义类别,难以扩展至开放世界。我们提出SeeGround,一种零样本3DVG框架,利用2D视觉语言模型(VLMs)避免3D特定训练。为弥合模态差距,引入混合输入格式,将查询对齐的渲染视图与空间丰富文本描述配对。框架包含两个核心组件:基于查询动态选择最优视角的视角适配模块,以及融合视觉与空间信号以提升定位精度的融合对齐模块。在ScanRefer和Nr3D上的大量评估表明,SeeGround显著优于现有零样本基线——分别提升7.7%和7.1%,甚至媲美全监督模型,在复杂条件下展现强大泛化能力。
原文摘要 · Abstract (English)
3D Visual Grounding (3DVG) seeks to locate target objects in 3D scenes using natural language descriptions, enabling downstream applications such as augmented reality and robotics. Existing approaches typically rely on labeled 3D data and predefined categories, limiting scalability to open-world settings. We present SeeGround, a zero-shot 3DVG framework that leverages 2D Vision-Language Models (VLMs) to bypass the need for 3D-specific training. To bridge the modality gap, we introduce a hybrid input format that pairs query-aligned rendered views with spatially enriched textual descriptions. Our framework incorporates two core components: a Perspective Adaptation Module that dynamically selects optimal viewpoints based on the query, and a Fusion Alignment Module that integrates visual and spatial signals to enhance localization precision. Extensive evaluations on ScanRefer and Nr3D confirm that SeeGround achieves substantial improvements over existing zero-shot baselines -- outperforming them by 7.7% and 7.1%, respectively -- and even rivals fully supervised alternatives, demonstrating strong generalization under challenging conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。