arXiv:2509.15871cs.CVcs.MM2025-09被引 5

用视图检索实现3D高斯点云的零样本视觉定位,无需每场景训练

Zero-Shot Visual Grounding in 3D Gaussians via View Retrieval

  • 将3D定位转为2D视图检索任务,利用多视角信息获取定位线索
  • 在多个数据集上达到当前最优性能,且无需场景特定训练
  • 适合需要快速部署、无标注数据的机器人等实际应用

3D视觉定位(3DVG)旨在根据文本提示定位3D场景中的物体,对机器人等应用至关重要。然而现有方法面临两大挑战:一是难以处理3D高斯溅射(3DGS)中空间纹理的隐式表示,导致必须进行每场景训练;二是通常需要大量标注数据才能有效训练。为此,我们提出基于视图检索的定位框架GVR,将3DVG转化为2D检索任务,通过物体级视图检索从多视角中收集定位线索,不仅避免了昂贵的3D标注过程,也消除了每场景训练的需求。大量实验表明,该方法在无需场景训练的情况下实现了最先进的视觉定位性能,为零样本3DVG研究提供了坚实基础。视频演示见https://github.com/leviome/GVR_demos。

原文摘要 · Abstract (English)

3D Visual Grounding (3DVG) aims to locate objects in 3D scenes based on text prompts, which is essential for applications such as robotics. However, existing 3DVG methods encounter two main challenges: first, they struggle to handle the implicit representation of spatial textures in 3D Gaussian Splatting (3DGS), making per-scene training indispensable; second, they typically require larges amounts of labeled data for effective training. To this end, we propose \underline{G}rounding via \underline{V}iew \underline{R}etrieval (GVR), a novel zero-shot visual grounding framework for 3DGS to transform 3DVG as a 2D retrieval task that leverages object-level view retrieval to collect grounding clues from multiple views, which not only avoids the costly process of 3D annotation, but also eliminates the need for per-scene training. Extensive experiments demonstrate that our method achieves state-of-the-art visual grounding performance while avoiding per-scene training, providing a solid foundation for zero-shot 3DVG research. Video demos can be found in https://github.com/leviome/GVR_demos.

3D视觉定位零样本3D高斯视图检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。