解决3D检测中新物体识别的语义不一致问题,提升开放词汇检测性能。
OV-SCAN: Semantically Consistent Alignment for Novel Object Discovery in Open-Vocabulary 3D Object Detection
- 通过精准3D标注和噪声对齐对过滤,实现跨模态语义一致对齐。
- 在nuScenes数据集上达到当前最优性能,有效识别训练外的新物体。
- 适合关注自动驾驶中开放词汇3D检测的研究者与工程师。
面向自动驾驶的开放词汇3D目标检测旨在识别点云场景中超出预定义标签集的新物体。现有方法通过将传统3D检测器与视觉语言模型(VLMs)结合,利用3D与2D特征间的跨模态对齐实现新物体的边界框回归与分类。然而,生成对应的3D与2D特征对时存在语义不一致,导致跨模态对齐效果不稳定。为此,本文提出OV-SCAN框架,通过两个核心策略克服该挑战:精确发现3D标注,并过滤由3D标注误差、遮挡或分辨率引起的低质量或损坏对齐对。在nuScenes数据集上的大量实验表明,OV-SCAN实现了当前最优性能。
原文摘要 · Abstract (English)
Open-vocabulary 3D object detection for autonomous driving aims to detect novel objects beyond the predefined training label sets in point cloud scenes. Existing approaches achieve this by connecting traditional 3D object detectors with vision-language models (VLMs) to regress 3D bounding boxes for novel objects and perform open-vocabulary classification through cross-modal alignment between 3D and 2D features. However, achieving robust cross-modal alignment remains a challenge due to semantic inconsistencies when generating corresponding 3D and 2D feature pairs. To overcome this challenge, we present OV-SCAN, an Open-Vocabulary 3D framework that enforces Semantically Consistent Alignment for Novel object discovery. OV-SCAN employs two core strategies: discovering precise 3D annotations and filtering out low-quality or corrupted alignment pairs (arising from 3D annotation, occlusion-induced, or resolution-induced noise). Extensive experiments on the nuScenes dataset demonstrate that OV-SCAN achieves state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。