CitySeg让无人机在城市级点云中实现无需视觉的开放词汇语义分割。
CitySeg: A 3D Open Vocabulary Semantic Segmentation Foundation Model in City-scale Scenarios
- 融合文本模态与分层图结构,提升模型对多源点云的感知能力
- 在9个封闭集基准上达最优,首次实现城市级点云零样本推理
- 适合自动驾驶、智慧城市等需要强泛化能力的三维感知场景
城市级点云语义分割是无人机感知系统的关键技术,可在不依赖视觉信息的情况下对3D点进行分类,实现全面的三维理解。然而,现有模型常受限于3D数据规模有限及数据集间的领域差异,导致泛化能力下降。为此,我们提出CitySeg,一个面向城市级点云语义分割的基础模型,通过引入文本模态实现开放词汇分割与零样本推理。为缓解多域数据分布不均问题,我们定制数据预处理规则,并设计局部-全局交叉注意力网络以增强点云网络在无人机场景下的感知能力。针对语义标签差异,提出分层分类策略:基于标注规则构建分层图,使用图编码器建模类别间层级关系。此外,采用两阶段训练策略并引入铰链损失,提升子类特征可分性。实验表明,CitySeg在9个封闭集基准上达到当前最优性能,显著优于现有方法。更重要的是,首次在城市级点云场景下实现了无需视觉信息的零样本泛化。
原文摘要 · Abstract (English)
Semantic segmentation of city-scale point clouds is a critical technology for Unmanned Aerial Vehicle (UAV) perception systems, enabling the classification of 3D points without relying on any visual information to achieve comprehensive 3D understanding. However, existing models are frequently constrained by the limited scale of 3D data and the domain gap between datasets, which lead to reduced generalization capability. To address these challenges, we propose CitySeg, a foundation model for city-scale point cloud semantic segmentation that incorporates text modality to achieve open vocabulary segmentation and zero-shot inference. Specifically, in order to mitigate the issue of non-uniform data distribution across multiple domains, we customize the data preprocessing rules, and propose a local-global cross-attention network to enhance the perception capabilities of point networks in UAV scenarios. To resolve semantic label discrepancies across datasets, we introduce a hierarchical classification strategy. A hierarchical graph established according to the data annotation rules consolidates the data labels, and the graph encoder is used to model the hierarchical relationships between categories. In addition, we propose a two-stage training strategy and employ hinge loss to increase the feature separability of subcategories. Experimental results demonstrate that the proposed CitySeg achieves state-of-the-art (SOTA) performance on nine closed-set benchmarks, significantly outperforming existing approaches. Moreover, for the first time, CitySeg enables zero-shot generalization in city-scale point cloud scenarios without relying on visual information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。