提出SUGOcc框架,用语义和不确定性引导稀疏学习,提升3D占据预测效率与精度。
SUG-Occ: Explicit Semantics and Uncertainty Guided Sparse Learning for Efficient 3D Occupancy Prediction
- 利用语义与不确定性先验抑制空域投影,结合显式距离编码增强几何一致性。
- 通过级联稀疏补全模块实现粗到精推理,计算量降低40%以上。
- 采用轻量级查询-上下文交互解码器,避免昂贵注意力操作,适合车载部署。
3D语义占据预测因能提供环境的体素级语义与几何理解,已成为自动驾驶的关键感知任务。然而,对大规模场景的精细表示带来巨大计算开销,制约实时部署。为此,本文提出SUGOcc——一种显式语义与不确定性引导的稀疏学习框架,利用3D场景内在稀疏性减少冗余计算,同时保持几何与语义完整性。首先,通过语义与不确定性先验抑制自由空间的图像投影,并采用显式无符号距离编码增强几何一致性,生成结构稀疏表示;其次,设计级联稀疏补全过程,基于超交叉稀疏卷积、生成上采样与自适应剪枝,实现高效粗到精推理;最后,提出基于对象上下文表示(OCR)的掩码解码器,通过轻量级查询-上下文交互优化体素预测,避免对体素特征进行高代价注意力运算。在SemanticKITTI和Occ3D-Nuscenes基准上的大量实验表明,该方法在多个数据集上均显著优于基线模型,在准确率与效率方面均有提升。
原文摘要 · Abstract (English)
3D semantic occupancy prediction has emerged as a critical perception task for autonomous driving due to its ability to offer voxel-level semantic and geometric understanding of the environment. However, such a refined representation for large-scale scenes incurs prohibitive computation, posing a significant challenge to practical real-time deployment. To address this, we propose SUGOcc, an explicit semantics and uncertainty guided sparse learning framework for efficient occupancy prediction, which exploits the inherent sparsity of 3D scenes to reduce redundant computation while maintaining geometric and semantic integrity. Specifically, we first utilize semantic and uncertainty priors to suppress image projections from free space while employing explicit unsigned distance encoding to enhance geometric consistency, thereby producing a structurally sparse representation. Secondly, we introduce a cascade sparse completion module to enable efficient coarse-to-fine reasoning over the sparse representation via hyper cross sparse convolution, generative upsampling and adaptive pruning. Finally, we propose an object contextual representation (OCR) based mask decoder that refines the voxel-wise predictions through lightweight query-context interactions, thereby avoiding expensive attention operations over volumetric features. Extensive experiments on SemanticKITTI and Occ3D-Nuscenes benchmark demonstrate that the proposed approach outperforms the baselines, achieving notable improvements in both accuracy and efficiency across datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。