融合激光雷达与摄像头,提升3D目标检测精度。
GAFusion: Adaptive Fusing LiDAR and Camera with Multiple Guidance for 3D Object Detection
- 用激光雷达引导生成深度和占据特征,增强多模态交互。
- 自适应融合变压器实现全局特征优化,提升检测性能。
- 适合自动驾驶中需要高精度3D感知的场景。
近年来基于鸟瞰图(BEV)视角的多模态3D目标检测方法取得了显著进展,但多数方法忽略了激光雷达与摄像头之间的互补交互与引导机制。本文提出一种新型多模态3D目标检测方法GAFusion,包含激光雷达引导的全局交互与自适应融合。具体地,引入稀疏深度引导(SDG)和激光雷达占据引导(LOG),生成富含深度信息的3D特征;设计激光雷达引导的自适应融合变压器(LGAFT),从全局角度自适应增强不同模态的BEV特征交互;同时,采用稀疏高度压缩下采样与多尺度双路径变压器(MSDPT),扩大各模态特征的感受野;最后引入时序融合模块,聚合前帧特征。在nuScenes测试集上,GAFusion达到73.6% mAP和74.9% NDS,性能达当前最优。
原文摘要 · Abstract (English)
Recent years have witnessed the remarkable progress of 3D multi-modality object detection methods based on the Bird's-Eye-View (BEV) perspective. However, most of them overlook the complementary interaction and guidance between LiDAR and camera. In this work, we propose a novel multi-modality 3D objection detection method, named GAFusion, with LiDAR-guided global interaction and adaptive fusion. Specifically, we introduce sparse depth guidance (SDG) and LiDAR occupancy guidance (LOG) to generate 3D features with sufficient depth information. In the following, LiDAR-guided adaptive fusion transformer (LGAFT) is developed to adaptively enhance the interaction of different modal BEV features from a global perspective. Meanwhile, additional downsampling with sparse height compression and multi-scale dual-path transformer (MSDPT) are designed to enlarge the receptive fields of different modal features. Finally, a temporal fusion module is introduced to aggregate features from previous frames. GAFusion achieves state-of-the-art 3D object detection results with 73.6$\%$ mAP and 74.9$\%$ NDS on the nuScenes test set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。