arXiv:2510.17686cs.CV2025-10NeurIPS被引 2

无需提示词,让3D检测模型在开放世界中发现未知物体。

Towards 3D Objectness Learning in an Open World

  • 用2D模型的语义先验和3D几何先验生成不依赖类别的候选框
  • 通过跨模态专家混合动态融合点云与图像特征,实现泛化检测
  • 在开放世界下比现有方法高16%性能,适合未知物体检测场景

近年来,3D目标检测与新类别检测取得显著进展,但对通用3D对象性学习的研究仍不足。本文聚焦开放世界3D对象性学习,旨在检测3D场景中所有物体,包括训练时未见的新类别。传统闭集3D检测器难以泛化到开放世界,而直接使用3D开放词汇模型则面临词汇扩展与语义重叠问题。为此,我们提出OP3Det:一种无类别提示、类无关的开放世界3D检测器,可在无需人工设计文本提示的情况下检测任意物体。利用2D基础模型的强大泛化能力,结合2D语义先验与3D几何先验生成类无关提案,拓展3D对象发现范围。通过跨模态专家混合机制,在点云与RGB图像间动态路由单模态与多模态特征,学习通用3D对象性。大量实验表明,OP3Det性能显著超越现有开放世界3D检测器,最高提升达16.0%(AR),相较闭集检测器提升13.5%。

原文摘要 · Abstract (English)

Recent advancements in 3D object detection and novel category detection have made significant progress, yet research on learning generalized 3D objectness remains insufficient. In this paper, we delve into learning open-world 3D objectness, which focuses on detecting all objects in a 3D scene, including novel objects unseen during training. Traditional closed-set 3D detectors struggle to generalize to open-world scenarios, while directly incorporating 3D open-vocabulary models for open-world ability struggles with vocabulary expansion and semantic overlap. To achieve generalized 3D object discovery, We propose OP3Det, a class-agnostic Open-World Prompt-free 3D Detector to detect any objects within 3D scenes without relying on hand-crafted text prompts. We introduce the strong generalization and zero-shot capabilities of 2D foundation models, utilizing both 2D semantic priors and 3D geometric priors for class-agnostic proposals to broaden 3D object discovery. Then, by integrating complementary information from point cloud and RGB image in the cross-modal mixture of experts, OP3Det dynamically routes uni-modal and multi-modal features to learn generalized 3D objectness. Extensive experiments demonstrate the extraordinary performance of OP3Det, which significantly surpasses existing open-world 3D detectors by up to 16.0% in AR and achieves a 13.5% improvement compared to closed-world 3D detectors.

3D检测开放世界零样本多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。