用相机和原始雷达数据融合,提升车载感知精度与效率。
A Resource Efficient Fusion Network for Object Detection in Bird's-Eye View using Camera and Raw Radar Data
- 直接使用雷达原始频谱,避免信号处理损失。
- 相机图像转鸟瞰极坐标域,与雷达特征融合检测目标。
- 在RADIal数据集上兼顾精度与计算开销,适合实际部署。
摄像头可感知车辆周围环境,而价格低廉的雷达传感器因能适应恶劣天气,在自动驾驶系统中广泛应用。然而,雷达点云稀疏,方位角与俯仰角分辨率低,缺乏场景的语义与结构信息,导致检测性能普遍较差。本文直接使用雷达的原始距离-多普勒(RD)谱,避免了雷达信号处理过程。在提出的综合图像处理流程中,独立处理摄像头图像:首先将图像转换至鸟瞰极坐标(BEV Polar)域,并通过相机编码器-解码器架构提取特征;随后将所得特征与从雷达解码器输入的径向-方位(RA)特征进行融合,实现目标检测。我们在RADIal数据集上评估了该融合策略,不仅对比了现有方法的检测精度,还考察了计算复杂度指标。
原文摘要 · Abstract (English)
Cameras can be used to perceive the environment around the vehicle, while affordable radar sensors are popular in autonomous driving systems as they can withstand adverse weather conditions unlike cameras. However, radar point clouds are sparser with low azimuth and elevation resolution that lack semantic and structural information of the scenes, resulting in generally lower radar detection performance. In this work, we directly use the raw range-Doppler (RD) spectrum of radar data, thus avoiding radar signal processing. We independently process camera images within the proposed comprehensive image processing pipeline. Specifically, first, we transform the camera images to Bird's-Eye View (BEV) Polar domain and extract the corresponding features with our camera encoder-decoder architecture. The resultant feature maps are fused with Range-Azimuth (RA) features, recovered from the RD spectrum input from the radar decoder to perform object detection. We evaluate our fusion strategy with other existing methods not only in terms of accuracy but also on computational complexity metrics on RADIal dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。