提出双引导融合方法,提升远距离小目标检测鲁棒性
DGFusion: Dual-guided Fusion for Robust Multi-Modal 3D Object Detection
- 基于难易度匹配构建双模态实例对,实现精准特征融合
- 在nuScenes数据集上提升1.0% mAP、0.8% NDS和1.3%平均召回率
- 特别适合处理遮挡、远距、小尺寸等困难目标检测场景
作为自动驾驶感知系统的关键任务,3D目标检测用于识别和追踪车辆、行人等关键物体。然而,远距离、小尺寸或被遮挡的目标(困难实例)的检测仍是难题,直接影响系统安全性。现有方法多采用单引导范式,未能考虑不同模态间困难实例的信息密度差异。本文提出DGFusion,基于双引导范式,融合点云引导图像与图像引导点云的优势。核心是难度感知实例配对器(DIPM),根据实例难易程度生成易/难实例对,双引导模块利用两类配对优势实现有效多模态特征融合。实验表明,DGFusion在nuScenes数据集上相比基线分别提升+1.0% mAP、+0.8% NDS和+1.3%平均召回率。在自车距离、尺寸、可见性及小样本训练场景下,对困难实例的检测性能均表现稳定提升。
原文摘要 · Abstract (English)
As a critical task in autonomous driving perception systems, 3D object detection is used to identify and track key objects, such as vehicles and pedestrians. However, detecting distant, small, or occluded objects (hard instances) remains a challenge, which directly compromises the safety of autonomous driving systems. We observe that existing multi-modal 3D object detection methods often follow a single-guided paradigm, failing to account for the differences in information density of hard instances between modalities. In this work, we propose DGFusion, based on the Dual-guided paradigm, which fully inherits the advantages of the Point-guide-Image paradigm and integrates the Image-guide-Point paradigm to address the limitations of the single paradigms. The core of DGFusion, the Difficulty-aware Instance Pair Matcher (DIPM), performs instance-level feature matching based on difficulty to generate easy and hard instance pairs, while the Dual-guided Modules exploit the advantages of both pair types to enable effective multi-modal feature fusion. Experimental results demonstrate that our DGFusion outperforms the baseline methods, with respective improvements of +1.0\% mAP, +0.8\% NDS, and +1.3\% average recall on nuScenes. Extensive experiments demonstrate consistent robustness gains for hard instance detection across ego-distance, size, visibility, and small-scale training scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。