arXiv:2502.06114cs.CV2025-02

用多路雷达特征融合提升稀疏输入下的3D目标检测精度

Enhanced 3D Object Detection via Diverse Feature Representations of 4D Radar Tensor

  • 通过多教师知识蒸馏融合不同预处理的4D雷达特征
  • 在极稀疏输入下实现AP_3D提升7.3%、AP_BEV提升9.5%
  • 输入数据量减少90倍,适合车载实时系统部署

近年来,车载四维(4D)雷达技术使原始4D雷达张量(4DRT)成为可能,其提供比传统点云更丰富的空间与多普勒信息。现有方法多依赖高度预处理的稀疏雷达数据,而直接利用原始4DRT面临计算开销大、可扩展性差的问题。为此,我们提出一种新型三维目标检测框架,最大化利用4DRT的同时保持高效。该方法引入多教师知识蒸馏(KD),多个教师模型分别在不同4DRT预处理方式生成的点云上训练,捕捉互补信号特征;这些教师表示通过专用聚合模块融合,并蒸馏至仅处理稀疏雷达输入的轻量级学生模型。在K-Radar数据集上的实验表明,当使用极稀疏输入时,本框架相较基线RTNH模型在AP_3D上提升7.3%,在AP_BEV上提升9.5%。此外,在显著降低输入数据量约90倍的情况下,性能接近密集输入基线,验证了方法的可扩展性与高效性。

原文摘要 · Abstract (English)

Recent advances in automotive four-dimensional (4D) Radar have enabled access to raw 4D Radar Tensor (4DRT), offering richer spatial and Doppler information than conventional point clouds. While most existing methods rely on heavily pre-processed, sparse Radar data, recent attempts to leverage raw 4DRT face high computational costs and limited scalability. To address these limitations, we propose a novel three-dimensional (3D) object detection framework that maximizes the utility of 4DRT while preserving efficiency. Our method introduces a multi-teacher knowledge distillation (KD), where multiple teacher models are trained on point clouds derived from diverse 4DRT pre-processing techniques, each capturing complementary signal characteristics. These teacher representations are fused via a dedicated aggregation module and distilled into a lightweight student model that operates solely on a sparse Radar input. Experimental results on the K-Radar dataset demonstrate that our framework achieves improvements of 7.3% in AP_3D and 9.5% in AP_BEV over the baseline RTNH model when using extremely sparse inputs. Furthermore, it attains comparable performance to denser-input baselines while significantly reducing the input data size by about 90 times, confirming the scalability and efficiency of our approach.

3D检测雷达感知知识蒸馏自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。