用2D视觉模型无标注实现大规模室外3D语义分割
Leveraging 2D-VLM for Label-Free 3D Segmentation in Large-Scale Outdoor Scene Understanding
- 通过虚拟相机将点云投影到2D图像,用语言提示驱动2D模型分割
- 多视角预测加权投票,实现无需3D标注的3D分割,精度接近有监督方法
- 支持任意文本查询的开放词汇识别,适合快速扩展新类别
本文提出一种新型3D语义分割方法,用于大规模点云数据,无需3D训练标注或配对的RGB图像。该方法利用虚拟相机将3D点云投影至2D图像,并通过自然语言提示引导的基础2D模型进行语义分割。3D分割结果通过多视角预测的加权投票聚合获得。所提方法优于现有无训练方法,在性能上接近有监督方法。此外,其支持开放词汇识别,用户可使用任意文本查询检测物体,突破传统监督方法的类别限制。
原文摘要 · Abstract (English)
This paper presents a novel 3D semantic segmentation method for large-scale point cloud data that does not require annotated 3D training data or paired RGB images. The proposed approach projects 3D point clouds onto 2D images using virtual cameras and performs semantic segmentation via a foundation 2D model guided by natural language prompts. 3D segmentation is achieved by aggregating predictions from multiple viewpoints through weighted voting. Our method outperforms existing training-free approaches and achieves segmentation accuracy comparable to supervised methods. Moreover, it supports open-vocabulary recognition, enabling users to detect objects using arbitrary text queries, thus overcoming the limitations of traditional supervised approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。