arXiv:2505.13817cs.CV2025-05

将实例与鸟瞰图特征融合,提升3D全景分割效率与精度

InstanceBEV: Unifying Instance and BEV Representation for 3D Panoptic Segmentation

  • 提出InstanceBEV,融合地图与目标中心建模优势
  • 8帧输入下达RayPQ 15.3、RayIoU 38.2,优于SparseOcc
  • 无需额外模块即可实现多任务协同,适合自动驾驶感知

基于鸟瞰图(BEV)的3D感知已成为端到端自动驾驶研究焦点。然而,现有BEV方法因特征空间过大,难以高效建模且阻碍全局注意力机制的有效集成。本文提出一种新策略InstanceBEV,协同地图中心与目标中心方法的优势。该方法在BEV特征中有效提取实例级特征,使全局注意力能在高度压缩的特征空间中实现,缓解了地图中心建模的效率问题。此外,本方法无需引入额外模块即可实现有效多任务学习。通过预测占用率,结合实例信息完成3D占用全景分割。在OCC3D-nuScenes数据集上的实验表明,仅使用8帧输入时,InstanceBEV达到RayPQ 15.3、RayIoU 38.2,相比SparseOcc分别提升9.3%和10.7%,验证了多任务协同的有效性。

原文摘要 · Abstract (English)

BEV-based 3D perception has emerged as a focal point of research in end-to-end autonomous driving. However, existing BEV approaches encounter significant challenges due to the large feature space, complicating efficient modeling and hindering effective integration of global attention mechanisms. We propose a novel modeling strategy, called InstanceBEV, that synergistically combines the strengths of both map-centric approaches and object-centric approaches. Our method effectively extracts instance-level features within the BEV features, facilitating the implementation of global attention modeling in a highly compressed feature space, thereby addressing the efficiency challenges inherent in map-centric global modeling. Furthermore, our approach enables effective multi-task learning without introducing additional module. We validate the efficiency and accuracy of the proposed model through predicting occupancy, achieving 3D occupancy panoptic segmentation by combining instance information. Experimental results on the OCC3D-nuScenes dataset demonstrate that InstanceBEV, utilizing only 8 frames, achieves a RayPQ of 15.3 and a RayIoU of 38.2. This surpasses SparseOcc's RayPQ by 9.3% and RayIoU by 10.7%, showcasing the effectiveness of multi-task synergy.

3D分割鸟瞰图多任务学习自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。