提出高效相机激光雷达融合3D语义占位方法,解决计算浪费与遮挡问题。
MR-Occ: Efficient Camera-LiDAR 3D Semantic Occupancy Prediction Using Hierarchical Multi-Resolution Voxel Representation
- 分层多分辨率体素优化关键区域特征,降低计算开销。
- 引入'被遮挡'类别,提升复杂场景下占位准确率。
- 通过可变形注意力融合图像与点云,适合自动驾驶感知任务。
精确的3D环境感知对自动驾驶至关重要。近期基于相机-激光雷达融合的3D语义占位方法虽提升了鲁棒性与准确性,但普遍存在计算资源均匀分配导致效率低下、遮挡处理不足影响精度的问题。本文提出MR-Occ,通过三个核心组件:分层体素特征精炼(HVFR)、多尺度占位解码器(MOD)和像素到体素融合网络(PVF-Net),有效应对上述挑战。HVFR增强关键体素特征,减少冗余计算;MOD引入‘被遮挡’类别,更精准建模传感器视域外区域;PVF-Net利用稠密化激光雷达特征,通过可变形注意力机制实现跨模态高效融合。大量实验表明,MR-Occ在nuScenes-Occupancy数据集上达到领先性能,IoU提升+5.2%,mIoU提升+5.3%,同时参数量与浮点运算量更低;在SemanticKITTI数据集上亦表现优异,验证了其在多种3D语义占位基准上的有效性与泛化能力。
原文摘要 · Abstract (English)
Accurate 3D perception is essential for understanding the environment in autonomous driving. Recent advancements in 3D semantic occupancy prediction have leveraged camera-LiDAR fusion to improve robustness and accuracy. However, current methods allocate computational resources uniformly across all voxels, leading to inefficiency, and they also fail to adequately address occlusions, resulting in reduced accuracy in challenging scenarios. We propose MR-Occ, a novel approach for camera-LiDAR fusion-based 3D semantic occupancy prediction, addressing these challenges through three key components: Hierarchical Voxel Feature Refinement (HVFR), Multi-scale Occupancy Decoder (MOD), and Pixel to Voxel Fusion Network (PVF-Net). HVFR improves performance by enhancing features for critical voxels, reducing computational cost. MOD introduces an `occluded' class to better handle regions obscured from sensor view, improving accuracy. PVF-Net leverages densified LiDAR features to effectively fuse camera and LiDAR data through a deformable attention mechanism. Extensive experiments demonstrate that MR-Occ achieves state-of-the-art performance on the nuScenes-Occupancy dataset, surpassing previous approaches by +5.2% in IoU and +5.3% in mIoU while using fewer parameters and FLOPs. Moreover, MR-Occ demonstrates superior performance on the SemanticKITTI dataset, further validating its effectiveness and generalizability across diverse 3D semantic occupancy benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。