无需标注和图像,直接从点云实现任意语言描述的语义分割。
OpenUrban3D: Annotation-Free Open-Vocabulary Semantic Segmentation of Large-Scale Urban Point Clouds
- 通过多视角渲染提取视觉-语言特征,生成鲁棒语义表示。
- 在SensatUrban和SUM上实现显著更高的准确率与跨场景泛化能力。
- 适合数字孪生、智慧城市等需要灵活识别新类别的应用。
开放词汇语义分割使模型能够识别并分割任意自然语言描述的物体,具备处理新类别、细粒度或功能定义类别的灵活性,对支持数字孪生、智慧城市建设与城市分析的大规模城市点云至关重要。然而,该能力在该领域仍基本未被探索。主要障碍在于大规模城市点云数据集通常缺乏高质量对齐的多视角影像,且现有三维分割流程在几何、尺度和外观差异较大的城市环境中泛化能力差。为此,我们提出OpenUrban3D,首个无需对齐多视图图像、预训练点云分割网络或人工标注的大规模城市场景3D开放词汇语义分割框架。方法通过多视角、多粒度渲染,提取掩码级视觉-语言特征,并进行样本均衡融合,再将特征蒸馏至3D主干模型。该设计支持零样本分割任意文本查询,同时保留语义丰富性与几何先验。在SensatUrban和SUM等大规模城市基准上的实验表明,OpenUrban3D在分割精度和跨场景泛化方面均显著优于现有方法,展现出作为灵活可扩展的3D城市场景理解方案的潜力。
原文摘要 · Abstract (English)
Open-vocabulary semantic segmentation enables models to recognize and segment objects from arbitrary natural language descriptions, offering the flexibility to handle novel, fine-grained, or functionally defined categories beyond fixed label sets. While this capability is crucial for large-scale urban point clouds that support applications such as digital twins, smart city management, and urban analytics, it remains largely unexplored in this domain. The main obstacles are the frequent absence of high-quality, well-aligned multi-view imagery in large-scale urban point cloud datasets and the poor generalization of existing three-dimensional (3D) segmentation pipelines across diverse urban environments with substantial variation in geometry, scale, and appearance. To address these challenges, we present OpenUrban3D, the first 3D open-vocabulary semantic segmentation framework for large-scale urban scenes that operates without aligned multi-view images, pre-trained point cloud segmentation networks, or manual annotations. Our approach generates robust semantic features directly from raw point clouds through multi-view, multi-granularity rendering, mask-level vision-language feature extraction, and sample-balanced fusion, followed by distillation into a 3D backbone model. This design enables zero-shot segmentation for arbitrary text queries while capturing both semantic richness and geometric priors. Extensive experiments on large-scale urban benchmarks, including SensatUrban and SUM, show that OpenUrban3D achieves significant improvements in both segmentation accuracy and cross-scene generalization over existing methods, demonstrating its potential as a flexible and scalable solution for 3D urban scene understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。