LXLv2提升雷达摄像头融合3D目标检测精度与速度
LXLv2: Enhanced LiDAR Excluded Lean 3D Object Detection with Fusion of 4D Radar and Camera
- 用雷达点云设计一多深度监督,提升深度一致性
- 引入注意力融合模块,检测精度提升1.8%(mAP)
- 适合自动驾驶中复杂场景下的实时目标感知
作为此前基于4D雷达-相机融合的3D目标检测最先进方法,LXL利用预测的图像深度分布图和雷达3D占用网格辅助基于采样的图像视图变换。然而,深度预测存在准确性和一致性不足的问题,且LXL采用拼接式融合方式影响了模型鲁棒性。本文提出LXLv2,通过改进克服上述局限并提升性能。具体地,针对雷达测量位置误差,设计基于雷达点的‘一对多’深度监督策略,并利用雷达回波强度(RCS)值调整监督区域,以实现物体级深度一致性;同时引入通道与空间注意力融合模块CSAFusion,增强特征适应性。在View-of-Delft和TJ4DRadSet数据集上的实验表明,所提LXLv2在检测精度、推理速度和鲁棒性方面均优于LXL,验证了模型有效性。
原文摘要 · Abstract (English)
As the previous state-of-the-art 4D radar-camera fusion-based 3D object detection method, LXL utilizes the predicted image depth distribution maps and radar 3D occupancy grids to assist the sampling-based image view transformation. However, the depth prediction lacks accuracy and consistency, and the concatenation-based fusion in LXL impedes the model robustness. In this work, we propose LXLv2, where modifications are made to overcome the limitations and improve the performance. Specifically, considering the position error in radar measurements, we devise a one-to-many depth supervision strategy via radar points, where the radar cross section (RCS) value is further exploited to adjust the supervision area for object-level depth consistency. Additionally, a channel and spatial attention-based fusion module named CSAFusion is introduced to improve feature adaptiveness. Experimental results on the View-of-Delft and TJ4DRadSet datasets show that the proposed LXLv2 can outperform LXL in detection accuracy, inference speed and robustness, demonstrating the effectiveness of the model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。