arXiv:2502.04377cs.CVcs.AI2025-02被引 48

提出新融合网络,让摄像头与激光雷达地图更准更对齐。

MapFusion: A Novel BEV Feature Fusion Network for Multi-modal Map Construction

  • 用跨模态交互模块解决图像与点云特征错位问题。
  • 在nuScenes数据集上,高清地图和鸟瞰图分割分别提升3.6%和6.2%。
  • 结构简洁易插拔,适合集成到现有自动驾驶地图系统中。

地图构建在自动驾驶系统中至关重要,需提供精确的静态环境信息。主流传感器包括摄像头和激光雷达,配置方式有纯摄像头、纯激光雷达或两者融合,通常融合方法表现最佳。然而,现有方法常忽视模态间交互,依赖简单融合策略,导致特征错位和信息丢失。为此,我们提出MapFusion,一种新型多模态鸟瞰图(BEV)特征融合方法。针对摄像头与激光雷达BEV特征间的语义错位问题,引入交叉模态交互变换(CIT)模块,通过自注意力机制实现两特征空间的交互,增强表示能力。同时提出双动态融合(DDF)模块,自适应选择各模态有价值信息,充分挖掘模态间互补性。此外,MapFusion设计简洁,可即插即用,易于融入现有流程。我们在两个地图构建任务——高精地图(HD map)与BEV地图分割上评估其性能,结果表明,在nuScenes数据集上,相比当前最优方法,分别实现3.6%和6.2%的绝对提升,验证了该方法的有效性与通用性。

原文摘要 · Abstract (English)

Map construction task plays a vital role in providing precise and comprehensive static environmental information essential for autonomous driving systems. Primary sensors include cameras and LiDAR, with configurations varying between camera-only, LiDAR-only, or camera-LiDAR fusion, based on cost-performance considerations. While fusion-based methods typically perform best, existing approaches often neglect modality interaction and rely on simple fusion strategies, which suffer from the problems of misalignment and information loss. To address these issues, we propose MapFusion, a novel multi-modal Bird's-Eye View (BEV) feature fusion method for map construction. Specifically, to solve the semantic misalignment problem between camera and LiDAR BEV features, we introduce the Cross-modal Interaction Transform (CIT) module, enabling interaction between two BEV feature spaces and enhancing feature representation through a self-attention mechanism. Additionally, we propose an effective Dual Dynamic Fusion (DDF) module to adaptively select valuable information from different modalities, which can take full advantage of the inherent information between different modalities. Moreover, MapFusion is designed to be simple and plug-and-play, easily integrated into existing pipelines. We evaluate MapFusion on two map construction tasks, including High-definition (HD) map and BEV map segmentation, to show its versatility and effectiveness. Compared with the state-of-the-art methods, MapFusion achieves 3.6% and 6.2% absolute improvements on the HD map construction and BEV map segmentation tasks on the nuScenes dataset, respectively, demonstrating the superiority of our approach.

地图构建多模态融合鸟瞰图自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。