arXiv:2411.16072cs.CV2024-11ICCV被引 21

用图像语义生成精准3D occupancy标签,提升开放词汇预测效果

Language Driven Occupancy Prediction

  • 通过图像到点云再到体素的语义传递,构建细粒度3D语言标签
  • 在Occ3D-nuScenes上优于现有零样本方法,显著减少人工标注
  • 通用框架适配多种模型结构,推动开放词汇占用预测发展

我们提出LOcc,一种高效且可泛化的开放词汇占用(OVO)预测框架。以往方法通常依赖图像特征作为中间媒介,通过粗粒度体素-文本对应或基于体素的模型-视图投影生成的噪声稀疏对应进行监督,导致监督信号不准确。为缓解此问题,我们设计了一种语义传递标注流程,生成密集且细粒度的3D语言占用真值。该流程有效挖掘图像中的语义信息,将文本标签从图像传递至激光雷达点云,并最终映射到体素,建立精确的体素-文本对应关系。通过将监督占用模型的原始预测头替换为二值占用状态的几何头和语言特征的语言头,LOcc利用生成的语言真值有效引导3D语言体积的学习。大量实验表明,我们的传递语义标注流程能生成更准确的伪标签,显著降低人工标注成本。此外,在多种架构上验证了LOcc,所有模型在Occ3D-nuScenes数据集上均持续优于当前最先进的零样本占用预测方法。

原文摘要 · Abstract (English)

We introduce LOcc, an effective and generalizable framework for open-vocabulary occupancy (OVO) prediction. Previous approaches typically supervise the networks through coarse voxel-to-text correspondences via image features as intermediates or noisy and sparse correspondences from voxel-based model-view projections. To alleviate the inaccurate supervision, we propose a semantic transitive labeling pipeline to generate dense and fine-grained 3D language occupancy ground truth. Our pipeline presents a feasible way to dig into the valuable semantic information of images, transferring text labels from images to LiDAR point clouds and ultimately to voxels, to establish precise voxel-to-text correspondences. By replacing the original prediction head of supervised occupancy models with a geometry head for binary occupancy states and a language head for language features, LOcc effectively uses the generated language ground truth to guide the learning of 3D language volume. Through extensive experiments, we demonstrate that our transitive semantic labeling pipeline can produce more accurate pseudo-labeled ground truth, diminishing labor-intensive human annotations. Additionally, we validate LOcc across various architectures, where all models consistently outperform state-of-the-art zero-shot occupancy prediction approaches on the Occ3D-nuScenes dataset.

3D理解占用预测多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。