arXiv:2509.15924cs.CV2025-09ICCV

无需训练,用2D模型实现稀疏视角下的开放词汇3D检测

Sparse Multiview Open-Vocabulary 3D Detection

  • 利用预训练2D模型生成2D检测,直接优化3D候选框
  • 在稀疏视角下性能超越现有方法,密集采样时也具竞争力
  • 适合无标注数据、计算资源有限的实时3D感知场景

三维场景理解对视觉与机器人系统至关重要。许多应用需要进行3D目标检测,即识别特定类别物体的位置与尺寸,通常以边界框表示。传统方法需针对固定类别训练,限制了实际应用。本文研究在挑战性但实用的稀疏视角设置下的开放词汇3D目标检测,仅依赖有限数量的带姿态RGB图像输入。提出的方法无需训练,仅使用现成的2D基础模型,避免昂贵的3D特征融合或专用3D学习。通过将2D检测提升至3D空间,并直接优化3D提案以实现跨视图特征度量一致性,充分挖掘2D模型在大规模训练数据上的优势。在标准基准测试中,该简单流程建立了一个强大基线,在密集采样场景下表现媲美最先进方法,而在稀疏视角下显著优于现有技术。

原文摘要 · Abstract (English)

The ability to interpret and comprehend a 3D scene is essential for many vision and robotics systems. In numerous applications, this involves 3D object detection, i.e.~identifying the location and dimensions of objects belonging to a specific category, typically represented as bounding boxes. This has traditionally been solved by training to detect a fixed set of categories, which limits its use. In this work, we investigate open-vocabulary 3D object detection in the challenging yet practical sparse-view setting, where only a limited number of posed RGB images are available as input. Our approach is training-free, relying on pre-trained, off-the-shelf 2D foundation models instead of employing computationally expensive 3D feature fusion or requiring 3D-specific learning. By lifting 2D detections and directly optimizing 3D proposals for featuremetric consistency across views, we fully leverage the extensive training data available in 2D compared to 3D. Through standard benchmarks, we demonstrate that this simple pipeline establishes a powerful baseline, performing competitively with state-of-the-art techniques in densely sampled scenarios while significantly outperforming them in the sparse-view setting.

3D检测开放词汇稀疏视角2D先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。