arXiv:2603.01007cs.CV2026-03中稿 · CVPR被引 1

用深度与区域信息提升自动驾驶3D占位预测精度

Dr.Occ: Depth- and Region-Guided 3D Occupancy from Surround-View Cameras for Autonomous Driving

  • 引入深度引导的2D到3D视角变换器,实现精准几何对齐
  • 设计区域自适应专家网络,缓解语义分布不均问题
  • 在nuScenes数据集上比基线提升7.43% mIoU,适合视觉感知研究

3D语义占位预测对自动驾驶感知至关重要,提供全面的几何场景理解与语义识别。然而,现有方法因缺乏像素级精确深度估计,在视图变换中存在几何错位,且面临严重空间类别不平衡问题,语义类别呈现强烈空间各向异性。为此,我们提出Dr. Occ,一种深度与区域引导的占位预测框架。具体而言,引入深度引导的2D到3D视图变换器(D²-VFormer),有效利用MoGe-2提供的高质量稠密深度线索构建可靠几何先验,实现体素特征的精确几何对齐。此外,受混合专家(MoE)框架启发,提出区域引导的专家变换器(R/R²-EFormer),自适应分配区域专属专家,聚焦不同空间区域,有效应对空间语义差异。两者互补:深度引导确保几何对齐,区域专家增强语义学习。在Occ3D--nuScenes基准上的实验表明,Dr. Occ在纯视觉设置下相比强基线BEVDet4D提升7.43% mIoU和3.09% IoU。

原文摘要 · Abstract (English)

3D semantic occupancy prediction is crucial for autonomous driving perception, offering comprehensive geometric scene understanding and semantic recognition. However, existing methods struggle with geometric misalignment in view transformation due to the lack of pixel-level accurate depth estimation, and severe spatial class imbalance where semantic categories exhibit strong spatial anisotropy. To address these challenges, we propose Dr. Occ, a depth- and region-guided occupancy prediction framework. Specifically, we introduce a depth-guided 2D-to-3D View Transformer (D$^2$-VFormer) that effectively leverages high-quality dense depth cues from MoGe-2 to construct reliable geometric priors, thereby enabling precise geometric alignment of voxel features. Moreover, inspired by the Mixture-of-Experts (MoE) framework, we propose a region-guided Expert Transformer (R/R$^2$-EFormer) that adaptively allocates region-specific experts to focus on different spatial regions, effectively addressing spatial semantic variations. Thus, the two components make complementary contributions: depth guidance ensures geometric alignment, while region experts enhance semantic learning. Experiments on the Occ3D--nuScenes benchmark demonstrate that Dr. Occ improves the strong baseline BEVDet4D by 7.43% mIoU and 3.09% IoU under the full vision-only setting.

3D占位自动驾驶深度估计多视角融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。