用视觉语言模型挖掘3D检测中的稀有物体,提升自动驾驶安全性能。
Concept-based Explainable Data Mining with VLM for 3D Detection
- 结合2D视觉语言模型与异常检测,自动发现驾驶场景中稀有但关键的物体。
- 仅用少量数据即显著提升对拖车、自行车等难检测类别的检测精度。
- 通过概念标注减少人工标注负担,适合自动驾驶数据高效构建场景。
稀有物体检测在自动驾驶系统中仍具挑战性,尤其当仅依赖点云数据时。尽管视觉语言模型(VLM)在图像理解方面表现强劲,其在3D目标检测中通过智能数据挖掘的潜力尚未充分探索。本文提出一种新型跨模态框架,利用2D VLM从驾驶场景中识别并挖掘稀有物体,从而提升3D检测性能。该方法融合目标检测、语义特征提取、降维及多维度异常检测,形成可解释的系统化流程,有效识别驾驶场景中语义上重要的稀有对象。通过结合孤立森林与t-SNE-based异常检测,并引入基于概念的过滤机制,框架能精准提取并标注如施工车辆、摩托车、障碍物等关键概念。这大幅降低标注成本,聚焦于最具价值的训练样本。在nuScenes数据集上的实验表明,这种概念引导的数据挖掘策略在仅使用极小部分训练数据的情况下,显著提升了3D检测模型性能,尤其在拖车和自行车等困难类别上优于同等规模的随机数据。该成果对安全关键型自动驾驶系统的高效数据集构建具有重要意义。
原文摘要 · Abstract (English)
Rare-object detection remains a challenging task in autonomous driving systems, particularly when relying solely on point cloud data. Although Vision-Language Models (VLMs) exhibit strong capabilities in image understanding, their potential to enhance 3D object detection through intelligent data mining has not been fully explored. This paper proposes a novel cross-modal framework that leverages 2D VLMs to identify and mine rare objects from driving scenes, thereby improving 3D object detection performance. Our approach synthesizes complementary techniques such as object detection, semantic feature extraction, dimensionality reduction, and multi-faceted outlier detection into a cohesive, explainable pipeline that systematically identifies rare but critical objects in driving scenes. By combining Isolation Forest and t-SNE-based outlier detection methods with concept-based filtering, the framework effectively identifies semantically meaningful rare objects. A key strength of this approach lies in its ability to extract and annotate targeted rare object concepts such as construction vehicles, motorcycles, and barriers. This substantially reduces the annotation burden and focuses only on the most valuable training samples. Experiments on the nuScenes dataset demonstrate that this concept-guided data mining strategy enhances the performance of 3D object detection models while utilizing only a fraction of the training data, with particularly notable improvements for challenging object categories such as trailers and bicycles compared with the same amount of random data. This finding has substantial implications for the efficient curation of datasets in safety-critical autonomous systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。