用2D大模型自动给3D点云打标签,保持跨领域一致性。
LeAP: Consistent multi-domain 3D labeling using Foundation Models
- 利用2D视觉大模型为3D点云生成任意类别的语义标签。
- 通过贝叶斯更新融合点标签到体素,提升时空一致性。
- 新领域适配后分割性能最高提升34.2 mIoU,适合无标注场景。
3D语义理解研究依赖数据集,虽获取未标注3D点云较容易,但手动标注成本高。近期2D视觉基础模型(VFMs)可实现开放集图像语义分割,可能助力自动化标注。然而,现有3D VFMs多基于2D模型改造,易引入标签不一致问题。本文提出LeAP(Label Any Pointcloud),利用2D VFMs自动为3D数据标注任意类别,同时保证标签一致性。通过贝叶斯更新将点级标签合并至体素,增强时空一致性;设计新型3D一致性网络(3D-CN)进一步利用3D信息优化标签质量。实验表明,该方法可在多种领域生成高质量3D语义标签,无需人工标注。在新领域适配中,模型语义分割的mIoU最高提升34.2。
原文摘要 · Abstract (English)
Availability of datasets is a strong driver for research on 3D semantic understanding, and whilst obtaining unlabeled 3D point cloud data is straightforward, manually annotating this data with semantic labels is time-consuming and costly. Recently, Vision Foundation Models (VFMs) enable open-set semantic segmentation on camera images, potentially aiding automatic labeling. However,VFMs for 3D data have been limited to adaptations of 2D models, which can introduce inconsistencies to 3D labels. This work introduces Label Any Pointcloud (LeAP), leveraging 2D VFMs to automatically label 3D data with any set of classes in any kind of application whilst ensuring label consistency. Using a Bayesian update, point labels are combined into voxels to improve spatio-temporal consistency. A novel 3D Consistency Network (3D-CN) exploits 3D information to further improve label quality. Through various experiments, we show that our method can generate high-quality 3D semantic labels across diverse fields without any manual labeling. Further, models adapted to new domains using our labels show up to a 34.2 mIoU increase in semantic segmentation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。