用小波分析融合雷达与摄像头数据,提升恶劣天气下3D目标检测性能
Wavelet-based Multi-View Fusion of 4D Radar Tensor and Camera for Robust 3D Object Detection

- 通过小波注意力机制与分视图特征提取,保留稀疏雷达信号关键信息
- 在雨雪条件下检测精度比现有方法高1.6%,全场景领先2.4%
- 适合自动驾驶中需全天候感知的多模态感知系统研发者
4D毫米波雷达因成本低、全天候鲁棒性强,广泛应用于自动驾驶与机器人感知。然而,基于点云的雷达表示在多阶段信号处理中存在信息损失,而直接使用原始4D雷达张量则计算开销巨大。为此,本文提出WRCFormer框架,通过解耦的多视图雷达表示,高效融合原始4D雷达立方体与相机图像。该方法包含两个核心组件:(1) 嵌入小波特征金字塔网络(FPN)的小波注意力模块,通过捕捉联合时空-频率特征,增强稀疏雷达信号与图像数据的表示能力,缓解信息丢失并保持计算效率;(2) 基于几何引导的渐进式融合机制,采用两阶段查询式融合策略,借助几何先验逐步对齐多视角雷达与视觉特征,实现无需大量计算开销的模态无关融合。在K-Radar基准测试中,WRCFormer在所有场景下超越现有最佳模型约2.4%,在雨夹雪条件下提升1.6%,展现出强抗恶劣天气能力。
原文摘要 · Abstract (English)
4D millimeter-wave (mmWave) radar has been widely adopted in autonomous driving and robot perception due to its low cost and all-weather robustness. However, point-cloud-based radar representations suffer from information loss due to multi-stage signal processing, while directly utilizing raw 4D radar tensors incurs prohibitive computational costs. To address these challenges, we propose WRCFormer, a novel 3D object detection framework that efficiently fuses raw 4D radar cubes with camera images via decoupled multi-view radar representations. Our approach introduces two key components: (1) A Wavelet Attention Module embedded in a wavelet-based Feature Pyramid Network (FPN), which enhances the representation of sparse radar signals and image data by capturing joint spatial-frequency features, thereby mitigating information loss while maintaining computational efficiency. (2) A Geometry-guided Progressive Fusion mechanism, a two-stage query-based fusion strategy that progressively aligns multi-view radar and visual features through geometric priors, enabling modality-agnostic and efficient integration without overwhelming computational overhead. Extensive experiments on the K-Radar benchmark show that WRCFormer achieves state-of-the-art performance, surpassing the best existing model by approximately 2.4% in all scenarios and 1.6% in sleet conditions, demonstrating strong robustness in adverse weather.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。