融合摄像头与雷达数据,提升复杂环境下的感知效率与精度。
REFNet++: Multi-Task Efficient Fusion of Camera and Radar Sensor Data in Bird's-Eye Polar View

- 将相机图像与雷达原始谱图映射到统一的鸟瞰极坐标视图
- 在RADIal数据集上实现车辆检测与自由空间分割性能领先
- 兼顾高精度与低计算开销,适合车载实时系统部署
摄像头能提供车辆周围环境的清晰视图,对环境感知至关重要;而价格低廉的雷达传感器因在多变天气下表现稳定而日益重要。然而,由于输出噪声大且分类能力弱,雷达需与其他传感器数据结合使用。为此,本文提出一种多任务高效融合方法,将雷达与摄像头数据对齐至统一的鸟瞰极坐标(BEV)域,兼顾精度与计算效率。模型以雷达的原始距离-多普勒(RD)谱和前视相机图像为输入,采用变分编码器-解码器结构,学习将前视图像转换为鸟瞰极坐标视图,同时雷达模块从RD数据中恢复角度信息,生成范围-方位(RA)特征。该对齐机制使双模态数据在兼容域中表示,从而实现鲁棒高效的传感器融合。我们在RADIal数据集上评估了车辆检测与自由空间分割任务,结果优于现有先进方法。
原文摘要 · Abstract (English)
A realistic view of the vehicle's surroundings is generally offered by camera sensors, which is crucial for environmental perception. Affordable radar sensors, on the other hand, are becoming invaluable due to their robustness in variable weather conditions. However, because of their noisy output and reduced classification capability, they work best when combined with other sensor data. Specifically, we address the challenge of multimodal sensor fusion by aligning radar and camera data in a unified domain, prioritizing not only accuracy, but also computational efficiency. Our work leverages the raw range-Doppler (RD) spectrum from radar and front-view camera images as inputs. To enable effective fusion, we employ a variational encoder-decoder architecture that learns the transformation of front-view camera data into the Bird's-Eye View (BEV) polar domain. Concurrently, a radar encoder-decoder learns to recover the angle information from the RD data that produce Range-Azimuth (RA) features. This alignment ensures that both modalities are represented in a compatible domain, facilitating robust and efficient sensor fusion. We evaluated our fusion strategy for vehicle detection and free space segmentation against state-of-the-art methods using the RADIal dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。