arXiv:2503.20614eess.IV2025-03ICRA被引 2

融合雷达与相机数据,提升复杂环境下的长距离3D目标检测鲁棒性。

SaViD: Spectravista Aesthetic Vision Integration for Robust and Discerning 3D Object Detection in Challenging Environments

  • 三阶段对齐机制融合稀疏激光雷达与高分辨率图像特征
  • 在Argoverse-2上实现9.87%的AP提升,Waymo数据集上mAPH提升2.39%
  • 对14类自然传感器退化具有强鲁棒性,性能提升超30%

激光雷达与相机的融合在自动驾驶短距离检测中表现优异,但在长距离场景下因数据稀疏性与分辨率差异面临挑战。此外,传感器退化进一步影响系统鲁棒性。本文提出SaViD框架,包含三个核心模块:全局记忆注意力网络(GMAN)增强图像特征提取;注意力稀疏记忆网络(ASMN)优化多模态特征融合;KNN连通图融合(KGF)实现空间信息完整融合。在长距离检测基准Argoverse-2上,平均精度(AP)提升9.87%,在Waymo Open Dataset(WOD)L2难度下mAPH提升2.39%。针对14种自然传感器退化,其相对错误率(RCE)在AV2上降低31.43%,在WOD上降低16.13%。代码已开源。

原文摘要 · Abstract (English)

The fusion of LiDAR and camera sensors has demonstrated significant effectiveness in achieving accurate detection for short-range tasks in autonomous driving. However, this fusion approach could face challenges when dealing with long-range detection scenarios due to disparity between sparsity of LiDAR and high-resolution camera data. Moreover, sensor corruption introduces complexities that affect the ability to maintain robustness, despite the growing adoption of sensor fusion in this domain. We present SaViD, a novel framework comprised of a three-stage fusion alignment mechanism designed to address long-range detection challenges in the presence of natural corruption. The SaViD framework consists of three key elements: the Global Memory Attention Network (GMAN), which enhances the extraction of image features through offering a deeper understanding of global patterns; the Attentional Sparse Memory Network (ASMN), which enhances the integration of LiDAR and image features; and the KNNnectivity Graph Fusion (KGF), which enables the entire fusion of spatial information. SaViD achieves superior performance on the long-range detection Argoverse-2 (AV2) dataset with a performance improvement of 9.87% in AP value and an improvement of 2.39% in mAPH for L2 difficulties on the Waymo Open dataset (WOD). Comprehensive experiments are carried out to showcase its robustness against 14 natural sensor corruptions. SaViD exhibits a robust performance improvement of 31.43% for AV2 and 16.13% for WOD in RCE value compared to other existing fusion-based methods while considering all the corruptions for both datasets. Our code is available at \href{https://github.com/sanjay-810/SAVID}

3D检测传感器融合鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。