arXiv:2505.03284cs.CVcs.RO2025-05被引 9

用柱面坐标融合多传感器数据,提升自动驾驶3D语义占位预测精度

OccCylindrical: Multi-Modal Fusion with Cylindrical Representation for 3D Semantic Occupancy Prediction

  • 在柱面坐标系下融合多模态感知特征,保留更多几何细节
  • 在nuScenes数据集上实现当前最佳性能,雨夜场景仍稳定有效
  • 适合关注自动驾驶环境感知与多传感器融合的研究者

自动驾驶车辆的安全运行高度依赖对周围环境的理解。3D语义占位预测任务将传感器周围的空域划分为体素,并为每个体素标注占据状态和语义信息。现有感知模型采用多传感器融合来完成该任务,但主要基于笛卡尔坐标系处理传感器信息,忽略了读数分布特性,导致细粒度几何细节丢失,性能下降。本文提出OccCylindrical,通过柱面坐标系融合并优化不同模态特征,有效保留了更精细的几何结构,从而提升预测性能。在nuScenes数据集上开展的大量实验,包括具有挑战性的雨天与夜间场景,验证了方法的有效性与领先性能。代码将公开于:https://github.com/DanielMing123/OccCylindrical

原文摘要 · Abstract (English)

The safe operation of autonomous vehicles (AVs) is highly dependent on their understanding of the surroundings. For this, the task of 3D semantic occupancy prediction divides the space around the sensors into voxels, and labels each voxel with both occupancy and semantic information. Recent perception models have used multisensor fusion to perform this task. However, existing multisensor fusion-based approaches focus mainly on using sensor information in the Cartesian coordinate system. This ignores the distribution of the sensor readings, leading to a loss of fine-grained details and performance degradation. In this paper, we propose OccCylindrical that merges and refines the different modality features under cylindrical coordinates. Our method preserves more fine-grained geometry detail that leads to better performance. Extensive experiments conducted on the nuScenes dataset, including challenging rainy and nighttime scenarios, confirm our approach's effectiveness and state-of-the-art performance. The code will be available at: https://github.com/DanielMing123/OccCylindrical

自动驾驶3D占位多模态融合柱面坐标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。