为激光雷达-相机融合设计轻量级图像主干,显著提升3D检测效率与精度。
DeGuNet: Depth-Guided Ultra-Compact Backbones for Efficient LiDAR-Camera 3D Detection

- 设计深度引导的稀疏特征提取机制,适配激光雷达点云的稀疏性。
- 在nuScenes上减少66.5%显存占用,推理提速1.16倍,mAP提升6.20点。
- 可直接替换现有模型主干,适合部署于资源受限的自动驾驶系统。
在自动驾驶感知中,激光雷达与摄像头的多模态融合已成为三维目标检测的主流范式。然而,当前框架严重依赖在二维语义任务上预训练的大规模视觉主干,导致参数冗余和结构不匹配,因二维先验难以应对鸟瞰图中激光雷达投影的极端稀疏性。为此,我们提出DeGuNet,一种专为深度引导表征学习设计的超轻量级、即插即用图像主干。通过引入稀疏感知特征提取机制,DeGuNet有效对齐多视角图像与非结构化激光雷达深度信息,严格避免无效区域污染。在nuScenes数据集上的大量实验表明,DeGuNet具备广泛的即插即用性与卓越的效率。集成至主流基线后,其彻底消除架构冗余,显存消耗最多降低66.5%,推理速度提升1.16倍,同时实现最高6.20绝对mAP提升,确立了参数高效多模态三维感知的新范式。
原文摘要 · Abstract (English)
In autonomous driving perception, the fusion of LiDAR and camera modalities has become the dominant paradigm for 3D object detection. However, current multi-modal frameworks heavily rely on massive visual backbones pretrained on 2D semantic tasks. This reliance introduces substantial parameter redundancy and a structural misalignment, as 2D priors are ill-equipped to handle the extreme sparsity of LiDAR projections required for Bird's-Eye-View geometry. To address this, we present DeGuNet, an ultra-compact and plug-and-play image backbone explicitly designed for depth-guided representation learning. By incorporating sparsity-aware feature extraction mechanisms, DeGuNet effectively aligns multi-view images with unstructured LiDAR depth while strictly preventing invalid-region contamination. Extensive experiments on the nuScenes dataset demonstrate DeGuNet's broad plug-and-play applicability and superior efficiency. When integrated into established baselines, it fundamentally eliminates architectural redundancy, reducing GPU memory consumption by up to 66.5% and achieving a 1.16x inference speedup. Concurrently, DeGuNet delivers up to a 6.20 absolute mAP gain, establishing a new paradigm for parameter-efficient multi-modal 3D perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。