用时序融合提升雷达相机3D检测,精度逼近激光雷达
RCTDistill: Cross-Modal Knowledge Distillation Framework for Radar-Camera 3D Object Detection with Temporal Fusion
- 设计三模块跨模态蒸馏框架,融合时序与空间信息
- 在nuScenes和VoD数据集上达到最新性能,推理速度26.2 FPS
- 适合做自动驾驶多传感器融合、边缘部署的工程师
雷达-相机融合方法已成为一种低成本的3D目标检测方案,但性能仍落后于基于激光雷达的方法。近期工作聚焦于时序融合与知识蒸馏(KD)策略以弥补差距,但尚未充分考虑物体运动或雷达与相机模态固有的不确定性。本文提出RCTDistill,一种基于时序融合的新型跨模态知识蒸馏框架,包含三个关键模块:距离-方位知识蒸馏(RAKD)、时序知识蒸馏(TKD)和区域解耦知识蒸馏(RDKD)。RAKD针对距离与方位方向的固有误差,实现从激光雷达特征到不准确鸟瞰图(BEV)表示的有效知识迁移;TKD通过对齐历史雷达-相机BEV特征与当前激光雷达表示,缓解动态物体引起的时序错位;RDKD通过蒸馏教师模型的关联知识,增强特征区分能力,使学生模型更好区分前景与背景。RCTDistill在nuScenes和View-of-Delft(VoD)数据集上均达到领先性能,且推理速度最快达26.2 FPS。
原文摘要 · Abstract (English)
Radar-camera fusion methods have emerged as a cost-effective approach for 3D object detection but still lag behind LiDAR-based methods in performance. Recent works have focused on employing temporal fusion and Knowledge Distillation (KD) strategies to overcome these limitations. However, existing approaches have not sufficiently accounted for uncertainties arising from object motion or sensor-specific errors inherent in radar and camera modalities. In this work, we propose RCTDistill, a novel cross-modal KD method based on temporal fusion, comprising three key modules: Range-Azimuth Knowledge Distillation (RAKD), Temporal Knowledge Distillation (TKD), and Region-Decoupled Knowledge Distillation (RDKD). RAKD is designed to consider the inherent errors in the range and azimuth directions, enabling effective knowledge transfer from LiDAR features to refine inaccurate BEV representations. TKD mitigates temporal misalignment caused by dynamic objects by aligning historical radar-camera BEV features with current LiDAR representations. RDKD enhances feature discrimination by distilling relational knowledge from the teacher model, allowing the student to differentiate foreground and background features. RCTDistill achieves state-of-the-art radar-camera fusion performance on both the nuScenes and View-of-Delft (VoD) datasets, with the fastest inference speed of 26.2 FPS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。