融合相机、激光雷达与雷达,提升车载鸟瞰图动态物体分割精度。
BEVMOSNet: Multimodal Fusion for BEV Moving Object Segmentation
- 采用可变形交叉注意力实现多模态传感器信息融合。
- 在nuScenes数据集上比单模态基线提升36.59%的分割交并比。
- 特别适用于夜间或雨天等恶劣环境下的自动驾驶系统。
在鸟瞰图(BEV)中准确理解动态物体的运动对确保自动驾驶车辆的可靠避障和顺畅路径规划至关重要。然而,相较于目标检测与分割任务,该问题研究较少,现有基于视觉的方法在夜间、低光照及雨天等恶劣条件下性能显著下降。相比之下,激光雷达和雷达在上述场景中表现稳定,且雷达能提供关键的速度信息。为此,本文提出BEVMOSNet,据我们所知是首个端到端融合相机、激光雷达与雷达的多模态方法,用于精确预测鸟瞰图中的运动物体。我们进一步深入分析,优化了基于可变形交叉注意力的跨传感器知识共享策略。在nuScenes数据集上的实验表明,相较于仅使用视觉的单模态基线BEV-MoSeg(Sigatapu et al., 2023),BEVMOSNet的总体交并比(IoU)提升了36.59%;相比扩展后的多模态方法SimpleBEV(Harley et al., 2022),也提升了2.35%,成为当前鸟瞰图运动分割的最先进方法。
原文摘要 · Abstract (English)
Accurate motion understanding of the dynamic objects within the scene in bird's-eye-view (BEV) is critical to ensure a reliable obstacle avoidance system and smooth path planning for autonomous vehicles. However, this task has received relatively limited exploration when compared to object detection and segmentation with only a few recent vision-based approaches presenting preliminary findings that significantly deteriorate in low-light, nighttime, and adverse weather conditions such as rain. Conversely, LiDAR and radar sensors remain almost unaffected in these scenarios, and radar provides key velocity information of the objects. Therefore, we introduce BEVMOSNet, to our knowledge, the first end-to-end multimodal fusion leveraging cameras, LiDAR, and radar to precisely predict the moving objects in BEV. In addition, we perform a deeper analysis to find out the optimal strategy for deformable cross-attention-guided sensor fusion for cross-sensor knowledge sharing in BEV. While evaluating BEVMOSNet on the nuScenes dataset, we show an overall improvement in IoU score of 36.59% compared to the vision-based unimodal baseline BEV-MoSeg (Sigatapu et al., 2023), and 2.35% compared to the multimodel SimpleBEV (Harley et al., 2022), extended for the motion segmentation task, establishing this method as the state-of-the-art in BEV motion segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。