arXiv:2411.16833cs.CV2024-11被引 25

用单张图片实现任意类别3D物体检测,突破传统模型类别限制。

Open Vocabulary Monocular 3D Object Detection

  • 融合预训练2D与3D视觉模型,减少对3D标注的依赖
  • 在未见类别上实现最佳零样本3D检测性能
  • 新评估指标缓解数据标签不全与语义模糊问题

我们提出并研究开放词汇单目3D目标检测这一新任务,旨在仅凭单张RGB图像,在度量3D空间中检测任意类别的物体。现有3D检测器要么依赖昂贵的激光雷达或多视角设置,要么局限于封闭词汇场景,类别有限,应用受限。我们识别出该设定下的两大挑战:一是3D边界框标注稀缺,影响模型泛化能力;为此,我们提出一个框架,有效整合预训练2D与3D视觉基础模型,降低对3D监督的依赖。二是现有数据集存在标签缺失和语义歧义(如桌子与书桌),妨碍可靠评估;为此,我们设计了一种新指标,能更准确衡量模型性能并缓解标注问题。我们的方法在未见类别的零样本3D检测以及已知类别的域内检测上均达到当前最优水平。我们希望本方法能成为强有力的基线,评估协议则为未来研究建立可靠基准。

原文摘要 · Abstract (English)

We propose and study open-vocabulary monocular 3D detection, a novel task that aims to detect objects of any categores in metric 3D space from a single RGB image. Existing 3D object detectors either rely on costly sensors such as LiDAR or multi-view setups, or remain confined to closed vocabularies settings with limited categories, restricting their applicability. We identify two key challenges in this new setting. First, the scarcity of 3D bounding box annotations limits the ability to train generalizable models. To reduce dependence on 3D supervision, we propose a framework that effectively integrates pretrained 2D and 3D vision foundation models. Second, missing labels and semantic ambiguities (\eg, table vs. desk) in existing datasets hinder reliable evaluation. To address this, we design a novel metric that captures model performance while mitigating annotation issues. Our approach achieves state-of-the-art results in zero-shot 3D detection of novel categories as well as in-domain detection on seen classes. We hope our method provides a strong baseline and our evaluation protocol establishes a reliable benchmark for future research.

3D检测开放词汇单目视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。