H3O高效预测3D占据网格,用多任务监督提升精度
H3O: Hyper-Efficient 3D Occupancy Prediction with Heterogeneous Supervision

- 采用轻量架构+异构监督,降低计算开销
- 在Occ3D-nuScenes上达到新高,优于当前最优方法
- 适合自动驾驶场景理解,兼顾效率与精度
3D占据预测作为全面理解三维场景的新范式,在自动驾驶规划中具有重要价值。现有方法大多计算成本高昂,依赖复杂的2D-3D注意力转换和3D特征处理。本文提出H3O,一种新型高效3D占据预测方法,其架构设计显著降低了计算开销。为缓解真实3D占据标签的模糊性,我们引入辅助任务进行补充监督:通过可微体渲染融合多相机深度估计、语义分割和表面法向估计,利用对应2D标签提供丰富异构监督信号。在Occ3D-nuScenes和SemanticKITTI基准上的大量实验表明,H3O性能优于当前最先进方法。
原文摘要 · Abstract (English)
3D occupancy prediction has recently emerged as a new paradigm for holistic 3D scene understanding and provides valuable information for downstream planning in autonomous driving. Most existing methods, however, are computationally expensive, requiring costly attention-based 2D-3D transformation and 3D feature processing. In this paper, we present a novel 3D occupancy prediction approach, H3O, which features highly efficient architecture designs that incur a significantly lower computational cost as compared to the current state-of-the-art methods. In addition, to compensate for the ambiguity in ground-truth 3D occupancy labels, we advocate leveraging auxiliary tasks to complement the direct 3D supervision. In particular, we integrate multi-camera depth estimation, semantic segmentation, and surface normal estimation via differentiable volume rendering, supervised by corresponding 2D labels that introduces rich and heterogeneous supervision signals. We conduct extensive experiments on the Occ3D-nuScenes and SemanticKITTI benchmarks that demonstrate the superiority of our proposed H3O.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。