用对比学习提升雷达相机特征对齐,显著改善3D目标检测性能。
Revisiting Radar Camera Alignment by Contrastive Learning for 3D Object Detection
- 基于对比学习设计双路径对齐模块,增强跨模态特征交互。
- 在nuScenes上实现4.3% NDS和8.4% mAP提升,达新基准。
- 针对雷达特征稀疏问题提出增强模块,适合自动驾驶感知场景。
基于雷达与相机融合的3D目标检测算法近期表现优异,为自动驾驶感知任务奠定了基础。现有方法多关注雷达与相机间因域差异导致的特征错位问题,但或忽略跨模态特征交互,或无法有效对齐同一空间位置的特征。为此,本文提出新型对齐模型RCAlign,设计基于对比学习的双路径对齐(DRA)模块,实现雷达与相机特征的对齐与融合。同时,针对雷达鸟瞰图(BEV)特征稀疏问题,引入雷达特征增强(RFE)模块,通过知识蒸馏损失提升雷达特征密度。实验表明,RCAlign在公开的nuScenes数据集上实现了雷达相机融合3D目标检测的新基准。相比最新方法RCBEVDet,其在实时3D检测中取得4.3% NDS和8.4% mAP的显著提升。
原文摘要 · Abstract (English)
Recently, 3D object detection algorithms based on radar and camera fusion have shown excellent performance, setting the stage for their application in autonomous driving perception tasks. Existing methods have focused on dealing with feature misalignment caused by the domain gap between radar and camera. However, existing methods either neglect inter-modal features interaction during alignment or fail to effectively align features at the same spatial location across modalities. To alleviate the above problems, we propose a new alignment model called Radar Camera Alignment (RCAlign). Specifically, we design a Dual-Route Alignment (DRA) module based on contrastive learning to align and fuse the features between radar and camera. Moreover, considering the sparsity of radar BEV features, a Radar Feature Enhancement (RFE) module is proposed to improve the densification of radar BEV features with the knowledge distillation loss. Experiments show RCAlign achieves a new state-of-the-art on the public nuScenes benchmark in radar camera fusion for 3D Object Detection. Furthermore, the RCAlign achieves a significant performance gain (4.3\% NDS and 8.4\% mAP) in real-time 3D detection compared to the latest state-of-the-art method (RCBEVDet).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。