arXiv:2510.06687cs.CVcs.AI2025-10

融合光场与激光雷达数据,提升复杂场景下语义分割精度。

Geometry-Aware Cross Modal Alignment for Light Field-LiDAR Semantic Segmentation

  • 通过特征补全模块解决点云与图像密度差异问题。
  • 深度感知模块提升对遮挡物体的分割准确率,mIoU提升2.38。
  • 首个集成光场与点云的多模态分割数据集,适合自动驾驶研究者。

语义分割是自动驾驶中场景理解的核心,但在遮挡等复杂条件下仍面临挑战。光场与激光雷达提供互补的视觉与空间信息,但其有效融合受限于视角多样性不足和模态差异。为此,我们首次提出一个整合光场数据与点云数据的多模态语义分割数据集,并基于此构建Mlpfseg网络,实现相机图像与激光雷达点云的联合分割。该网络包含特征补全模块,通过对点云特征图进行差异重建,缓解点云与图像像素间的密度不匹配;以及深度感知模块,通过增强注意力以提升对遮挡区域的感知能力。实验表明,该方法在图像单模态基础上提升1.71 mIoU,点云单模态基础上提升2.38 mIoU,验证了其有效性。

原文摘要 · Abstract (English)

Semantic segmentation serves as a cornerstone of scene understanding in autonomous driving but continues to face significant challenges under complex conditions such as occlusion. Light field and LiDAR modalities provide complementary visual and spatial cues that are beneficial for robust perception; however, their effective integration is hindered by limited viewpoint diversity and inherent modality discrepancies. To address these challenges, the first multimodal semantic segmentation dataset integrating light field data and point cloud data is proposed. Based on this dataset, we proposed a multi-modal light field point-cloud fusion segmentation network(Mlpfseg), incorporating feature completion and depth perception to segment both camera images and LiDAR point clouds simultaneously. The feature completion module addresses the density mismatch between point clouds and image pixels by performing differential reconstruction of point-cloud feature maps, enhancing the fusion of these modalities. The depth perception module improves the segmentation of occluded objects by reinforcing attention scores for better occlusion awareness. Our method outperforms image-only segmentation by 1.71 Mean Intersection over Union(mIoU) and point cloud-only segmentation by 2.38 mIoU, demonstrating its effectiveness.

语义分割多模态融合自动驾驶点云处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。