用4D雷达与图像融合提升自动驾驶感知鲁棒性
RadarXFormer: Robust Object Detection via Cross-Dimension Fusion of 4D Radar Spectra and Images for Autonomous Driving
- 直接处理原始雷达谱,构建高效3D表示减少数据量
- 跨维度融合3D雷达立方体与2D图像特征,提升检测精度
- 在恶劣条件下仍保持实时推理,适合智能驾驶应用
可靠的感知对自动驾驶系统在复杂交通条件下的安全运行至关重要。然而,基于摄像头和激光雷达的感知系统在恶劣天气和光照条件下性能下降,限制了其在智能交通系统中的广泛应用。雷达-视觉融合通过结合毫米波雷达的环境鲁棒性和低成本优势与摄像头的丰富语义信息,提供了一种有前景的替代方案。尽管传统3D雷达测量缺乏高度分辨率且数据稀疏,新兴的4D毫米波雷达引入了高程信息,但也带来了信号噪声大、数据量高等挑战。为此,本文提出RadarXFormer,一种3D目标检测框架,实现4D雷达谱与RGB图像的高效跨模态融合。该方法不依赖稀疏雷达点云,而是直接利用原始雷达谱,构建高效的3D表示,在降低数据量的同时保留完整的3D空间信息。'X'代表提出的跨维度(3D-2D)融合机制,多尺度3D球形雷达特征立方体与互补的2D图像特征图进行融合。在K-Radar数据集上的实验表明,该方法在复杂条件下显著提升了检测准确率与鲁棒性,同时保持实时推理能力。
原文摘要 · Abstract (English)
Reliable perception is essential for autonomous driving systems to operate safely under diverse real-world traffic conditions. However, camera- and LiDAR-based perception systems suffer from performance degradation under adverse weather and lighting conditions, limiting their robustness and large-scale deployment in intelligent transportation systems. Radar-vision fusion provides a promising alternative by combining the environmental robustness and cost efficiency of millimeter-wave (mmWave) radar with the rich semantic information captured by cameras. Nevertheless, conventional 3D radar measurements lack height resolution and remain highly sparse, while emerging 4D mmWave radar introduces elevation information but also brings challenges such as signal noise and large data volume. To address these issues, this paper proposes RadarXFormer, a 3D object detection framework that enables efficient cross-modal fusion between 4D radar spectra and RGB images. Instead of relying on sparse radar point clouds, RadarXFormer directly leverages raw radar spectra and constructs an efficient 3D representation that reduces data volume while preserving complete 3D spatial information. The "X" highlights the proposed cross-dimension (3D-2D) fusion mechanism, in which multi-scale 3D spherical radar feature cubes are fused with complementary 2D image feature maps. Experiments on the K-Radar dataset demonstrate improved detection accuracy and robustness under challenging conditions while maintaining real-time inference capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。