arXiv:2602.03361cs.CV2026-02ACL被引 1

仅用多视角图像实现零样本3D物体定位,无需几何标注

Z3D: Zero-Shot 3D Visual Grounding from Images

  • 基于多视角图像构建通用定位流程,可选使用相机位姿和深度图
  • 在ScanRefer和Nr3D上达到零样本方法最优性能
  • 结合先进提示分割与实例分割提升定位精度,适合无标注场景

3D视觉定位(3DVG)旨在根据自然语言查询定位3D场景中的物体。本文探索仅从多视角图像出发的零样本3DVG,无需任何几何监督或物体先验。提出Z3D,一种灵活适配多视角图像的通用定位框架,可选融合相机位姿和深度图。识别出先前零样本方法的关键瓶颈,并通过(i)采用前沿零样本3D实例分割生成高质量3D边界框提案,以及(ii)基于提示的高级推理机制,充分挖掘现代视觉语言模型(VLMs)的能力来解决。在ScanRefer和Nr3D基准上的大量实验表明,该方法在零样本3DVG中达到领先水平。代码已开源:https://github.com/col14m/z3d。

原文摘要 · Abstract (English)

3D visual grounding (3DVG) aims to localize objects in a 3D scene based on natural language queries. In this work, we explore zero-shot 3DVG from multi-view images alone, without requiring any geometric supervision or object priors. We introduce Z3D, a universal grounding pipeline that flexibly operates on multi-view images while optionally incorporating camera poses and depth maps. We identify key bottlenecks in prior zero-shot methods causing significant performance degradation and address them with (i) a state-of-the-art zero-shot 3D instance segmentation method to generate high-quality 3D bounding box proposals and (ii) advanced reasoning via prompt-based segmentation, which utilizes full capabilities of modern VLMs. Extensive experiments on the ScanRefer and Nr3D benchmarks demonstrate that our approach achieves state-of-the-art performance among zero-shot methods. Code is available at https://github.com/col14m/z3d .

3D定位零样本多视图视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。