arXiv:2507.08903cs.ROcs.CV2025-07中稿 · ITSC'25

用路边设备融合摄像头与激光雷达数据,提升复杂路口高清语义地图精度。

Multimodal HD Mapping for Intersections by Intelligent Roadside Units

  • 通过路边单元融合相机与激光雷达数据,解决车辆视角受限问题。
  • 在自建数据集上,语义分割mIoU比仅用图像高4%,比仅用点云高18%。
  • 适合研究智能交通基础设施与自动驾驶协同的学者和工程师。

复杂路口的高精语义地图构建对传统车载方法构成挑战,主要源于遮挡和视角局限。本文提出一种基于路侧智能单元(IRUs)的新型相机-LiDAR融合框架,并构建了RS-seq数据集,该数据集通过系统性增强和标注V2X-Seq数据集获得。RS-seq包含从路侧设备采集的精确标注图像与点云数据,以及七个路口的矢量化地图,涵盖车道线、人行横道、停止线等详细特征。该数据集支持对路侧数据在高精地图生成中跨模态互补性的系统研究。所提融合框架采用两阶段流程:先进行模态特异性特征提取,再实现跨模态语义融合,充分利用相机高分辨率纹理与激光雷达精准几何信息。在RS-seq数据集上的定量评估表明,多模态方法持续优于单模态基线。具体而言,相比仅使用图像的基线,多模态方法在语义分割上将平均交并比(mIoU)提升4%;相比仅使用点云的基线,提升18%。本研究为基于路侧单元的高精语义地图构建建立了基准方法,并提供了未来基础设施辅助自动驾驶研究的宝贵数据集。

原文摘要 · Abstract (English)

High-definition (HD) semantic mapping of complex intersections poses significant challenges for traditional vehicle-based approaches due to occlusions and limited perspectives. This paper introduces a novel camera-LiDAR fusion framework that leverages elevated intelligent roadside units (IRUs). Additionally, we present RS-seq, a comprehensive dataset developed through the systematic enhancement and annotation of the V2X-Seq dataset. RS-seq includes precisely labelled camera imagery and LiDAR point clouds collected from roadside installations, along with vectorized maps for seven intersections annotated with detailed features such as lane dividers, pedestrian crossings, and stop lines. This dataset facilitates the systematic investigation of cross-modal complementarity for HD map generation using IRU data. The proposed fusion framework employs a two-stage process that integrates modality-specific feature extraction and cross-modal semantic integration, capitalizing on camera high-resolution texture and precise geometric data from LiDAR. Quantitative evaluations using the RS-seq dataset demonstrate that our multimodal approach consistently surpasses unimodal methods. Specifically, compared to unimodal baselines evaluated on the RS-seq dataset, the multimodal approach improves the mean Intersection-over-Union (mIoU) for semantic segmentation by 4\% over the image-only results and 18\% over the point cloud-only results. This study establishes a baseline methodology for IRU-based HD semantic mapping and provides a valuable dataset for future research in infrastructure-assisted autonomous driving systems.

高精地图多模态融合路侧单元自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。