arXiv:2506.07002cs.CV2025-06被引 1

用双表示法提升3D占据预测精度与效率

BePo: Dual Representation for 3D Occupancy Prediction

  • 融合俯视图与稀疏点云的双表示结构
  • 在多个基准上达到领先性能,推理成本低
  • 适合自动驾驶场景中细粒度3D感知任务

3D占据预测对自动驾驶中的精细3D几何与语义理解至关重要。现有方法计算开销大,依赖密集3D特征体和交叉注意力来聚合信息。更高效的方法采用鸟瞰图(BEV)或稀疏点云作为场景表示,显著降低运行时间。但BEV难以表征小物体,尤其在投影到地面平面后特征稀疏;而稀疏点云虽能建模各类尺寸物体,却在捕捉平面或大物体时效率低下。为此,本文提出BePo,采用BEV与稀疏点云的双重表示。稀疏点云分支学习到的3D信息通过交叉注意力传递给BEV分支,向其注入难例物体的学习信号。两分支输出融合生成最终3D占据预测。在Occ3D-nuScenes、Occ3D-Waymo和Occ-ScanNet等挑战性基准上进行了大量实验,验证了BePo的优越性。此外,即使与最新高效方法相比,BePo仍保持较低推理成本。

原文摘要 · Abstract (English)

3D occupancy infers fine-grained 3D geometry and semantics which is critical for autonomous driving. Most existing approaches carry high compute costs, requiring dense 3D feature volume and cross-attention to effectively aggregate information. More efficient methods adopt Bird's Eye View (BEV) or sparse points as scene representation leading to much reduced runtime. However, BEV struggles with small objects that often have very limited feature representation especially after being projected to the ground plane. Sparse points on the other and, can model objects of various sizes in 3D space, but is inefficient at capturing flat surfaces or large objects. To address these shortcomings, we present BePo, which features a dual representation of BEV and sparse points. The 3D information learned in the sparse points branch is shared with the BEV stream via cross-attention, which injects learning signals of difficult objects on the BEV plane. The outputs of both branches are then fused to generate the final 3D occupancy predictions. Extensive experiments on a suite of challenging benchmarks including Occ3D-nuScenes, Occ3D-Waymo and Occ-ScanNet demonstrate the superiority of our proposed BePo. In addition, BePo carries low inference cost even when compared to latest efficient methods.

3D占据自动驾驶双表示高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。