提出新融合方法,提升激光雷达与摄像头3D目标检测精度
InsFusion: Rethink Instance-level LiDAR-Camera Fusion for 3D Object Detection

- 从原始和融合特征中提取候选框,用以查询原始特征
- 在nuScenes上达到新最好效果,兼容多种主流方法
- 适合自动驾驶、智能交通领域研究者参考
基于多视角摄像头与激光雷达的三维目标检测是自动驾驶与智慧交通的关键技术。然而,在基础特征提取、透视变换和特征融合过程中,噪声与误差会逐步累积。为解决此问题,本文提出InsFusion:从原始特征与融合特征中分别提取候选框,并利用这些候选框查询原始特征,从而缓解累积误差的影响。此外,通过在原始特征上引入注意力机制,进一步抑制误差传播。在nuScenes数据集上的实验表明,InsFusion可兼容多种先进基线方法,并在3D目标检测任务中实现新的最佳性能。
原文摘要 · Abstract (English)
Three-dimensional Object Detection from multi-view cameras and LiDAR is a crucial component for autonomous driving and smart transportation. However, in the process of basic feature extraction, perspective transformation, and feature fusion, noise and error will gradually accumulate. To address this issue, we propose InsFusion, which can extract proposals from both raw and fused features and utilizes these proposals to query the raw features, thereby mitigating the impact of accumulated errors. Additionally, by incorporating attention mechanisms applied to the raw features, it thereby mitigates the impact of accumulated errors. Experiments on the nuScenes dataset demonstrate that InsFusion is compatible with various advanced baseline methods and delivers new state-of-the-art performance for 3D object detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。