arXiv:2504.15888cs.CV2025-04被引 5

融合激光雷达与摄像头,提升复杂环境3D语义占位感知精度

MS-Occ: Multi-Stage LiDAR-Camera Fusion for 3D Semantic Occupancy Prediction

  • 分阶段融合:中层用几何增强图像,深层用自适应融合优化语义
  • nuScenes上达32.1% IoU,SemanticKITTI上mIoU达24.08%,刷新纪录
  • 对小物体感知显著提升,适合自动驾驶等安全关键场景

精确的3D语义占位感知对于复杂环境中多样且不规则物体的自动驾驶至关重要。视觉主导方法存在几何误差,而基于激光雷达的方法则缺乏丰富语义信息。为此,本文提出MS-Occ——一种包含中层与深层融合的多阶段激光雷达-相机融合框架,通过层级跨模态融合整合激光雷达的几何保真度与相机的语义丰富性。该框架在两个关键阶段引入创新:(1) 中层特征融合中,Gaussian-Geo模块利用高斯核渲染稀疏激光雷达深度图,为2D图像特征注入密集几何先验;Semantic-Aware模块通过可变形交叉注意力为激光雷达体素补充语义上下文;(2) 深层体素融合中,自适应融合(AF)模块动态平衡多模态体素特征,高置信度体素融合(HCCVF)模块通过基于自注意力的精炼解决语义不一致问题。在两个大规模基准测试上实验表明性能达到新高度:在nuScenes-OpenOccupancy上,交并比(IoU)达32.1%,均交并比(mIoU)为25.3%,分别优于现有最优方法0.7%和2.4%;在SemanticKITTI上,新取得24.08%的mIoU,充分验证其泛化能力。消融实验进一步证实各模块有效性,尤其在小物体感知上提升显著,凸显其在安全关键自动驾驶场景中的实用价值。

原文摘要 · Abstract (English)

Accurate 3D semantic occupancy perception is essential for autonomous driving in complex environments with diverse and irregular objects. While vision-centric methods suffer from geometric inaccuracies, LiDAR-based approaches often lack rich semantic information. To address these limitations, MS-Occ, a novel multi-stage LiDAR-camera fusion framework which includes middle-stage fusion and late-stage fusion, is proposed, integrating LiDAR's geometric fidelity with camera-based semantic richness via hierarchical cross-modal fusion. The framework introduces innovations at two critical stages: (1) In the middle-stage feature fusion, the Gaussian-Geo module leverages Gaussian kernel rendering on sparse LiDAR depth maps to enhance 2D image features with dense geometric priors, and the Semantic-Aware module enriches LiDAR voxels with semantic context via deformable cross-attention; (2) In the late-stage voxel fusion, the Adaptive Fusion (AF) module dynamically balances voxel features across modalities, while the High Classification Confidence Voxel Fusion (HCCVF) module resolves semantic inconsistencies using self-attention-based refinement. Experiments on two large-scale benchmarks demonstrate state-of-the-art performance. On nuScenes-OpenOccupancy, MS-Occ achieves an Intersection over Union (IoU) of 32.1% and a mean IoU (mIoU) of 25.3%, surpassing the state-of-the-art by +0.7% IoU and +2.4% mIoU. Furthermore, on the SemanticKITTI benchmark, our method achieves a new state-of-the-art mIoU of 24.08%, robustly validating its generalization capabilities.Ablation studies further confirm the effectiveness of each individual module, highlighting substantial improvements in the perception of small objects and reinforcing the practical value of MS-Occ for safety-critical autonomous driving scenarios.

3D占位多模态融合自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。