arXiv:2505.20951cs.CV2025-05

用深度感知与语义辅助,提升摄像头3D语义占位预测精度

DSOcc: Leveraging Depth Awareness and Semantic Aid to Boost Camera-Based 3D Semantic Occupancy Prediction

  • 联合推断占位状态与类别,用深度信心值加权图像特征
  • 在SemanticKITTI上超越现有摄像头方法,达最优性能
  • 适合自动驾驶场景感知,尤其关注低成本视觉方案

基于摄像头的3D语义占位预测为自动驾驶提供了高效且低成本的环境感知方案。然而,现有方法依赖显式占位状态推断,导致大量误分配特征;同时样本不足限制了占位类别学习。为此,本文提出利用深度感知与语义辅助提升摄像头3D语义占位预测(DSOcc)。我们联合进行占位状态与类别推断:通过非学习方法计算软占位置信度,并与图像特征相乘,使体素具备深度感知能力,实现自适应隐式占位状态推断。不增强特征学习,而是直接使用预训练的图像语义分割模型,融合多帧及其占位概率以辅助类别推断,提升鲁棒性。实验表明,DSOcc在SemanticKITTI上达到摄像头方法中的最先进性能,在SSCBench-KITTI-360和Occ3D-nuScenes上也表现竞争力。代码将开源。

原文摘要 · Abstract (English)

Camera-based 3D semantic occupancy prediction offers an efficient and cost-effective solution for perceiving surrounding scenes in autonomous driving. However, existing works rely on explicit occupancy state inference, leading to numerous incorrect feature assignments, and insufficient samples restrict the learning of occupancy class inference. To address these challenges, we propose leveraging \textbf{D}epth awareness and \textbf{S}emantic aid to boost camera-based 3D semantic \textbf{Occ}upancy prediction (\textbf{DSOcc}). We jointly perform occupancy state and occupancy class inference, where soft occupancy confidence is calculated by non-learning method and multiplied with image features to make voxels aware of depth, enabling adaptive implicit occupancy state inference. Instead of enhancing feature learning, we directly utilize well-trained image semantic segmentation and fuse multiple frames with their occupancy probabilities to aid occupancy class inference, thereby enhancing robustness. Experimental results demonstrate that DSOcc achieves state-of-the-art performance on the SemanticKITTI dataset among camera-based methods and achieves competitive performance on the SSCBench-KITTI-360 and Occ3D-nuScenes datasets. Code will be released on github.

3D占位语义感知自动驾驶视觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。