MinkOcc用少量人工标注实现高效3D语义占位预测,大幅降低标注成本。
MinkOcc: Towards real-time label-efficient semantic occupancy prediction
- 采用两阶段半监督训练,先用少量3D标注启动,再用视觉大模型生成的图像和点云标签继续训练。
- 标注依赖减少90%,在自动驾驶场景中保持高精度,支持实时推理。
- 融合相机与激光雷达数据,基于稀疏卷积网络,适合实际部署。
构建3D语义占位预测模型通常依赖密集的3D标注进行监督学习,这一过程耗时耗力。为解决该问题,本文提出MinkOcc,一种面向摄像头与激光雷达的多模态3D语义占位预测框架,采用两阶段半监督训练:首先利用少量显式3D标注启动训练;随后通过更易标注的累积激光雷达扫描与图像——其语义标签由视觉基础模型生成——持续提供监督信号。MinkOcc有效利用丰富的传感器监督信号,在保持竞争性精度的同时,将人工标注依赖降低90%。此外,模型通过早期融合整合激光雷达与相机信息,并采用稀疏卷积网络实现实时预测。凭借在标注与计算上的双重效率,本工作旨在推动MinkOcc超越受控数据集,实现自动驾驶中3D语义占位预测的广泛落地。
原文摘要 · Abstract (English)
Developing 3D semantic occupancy prediction models often relies on dense 3D annotations for supervised learning, a process that is both labor and resource-intensive, underscoring the need for label-efficient or even label-free approaches. To address this, we introduce MinkOcc, a multi-modal 3D semantic occupancy prediction framework for cameras and LiDARs that proposes a two-step semi-supervised training procedure. Here, a small dataset of explicitly 3D annotations warm-starts the training process; then, the supervision is continued by simpler-to-annotate accumulated LiDAR sweeps and images -- semantically labelled through vision foundational models. MinkOcc effectively utilizes these sensor-rich supervisory cues and reduces reliance on manual labeling by 90\% while maintaining competitive accuracy. In addition, the proposed model incorporates information from LiDAR and camera data through early fusion and leverages sparse convolution networks for real-time prediction. With its efficiency in both supervision and computation, we aim to extend MinkOcc beyond curated datasets, enabling broader real-world deployment of 3D semantic occupancy prediction in autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。