为城市级3D点云设计自监督学习方法,提升语义分割性能。
Polis: 3D Self-Supervision at City Scale

- 提出Polis框架,融合几何不变性与正则化损失,适配城市尺度点云
- 在14个城市数据集上达到23.8%平均mIoU,显著优于现有方法
- 适用于城市空间分析与自动驾驶,但对地面局部场景效果下降
从城市级3D模型中获取可靠的语义表示对城市分析、基础设施监测、自动驾驶和遗产保护日益重要。然而,通过航空测绘获取的大范围城市场景与用于预训练多数3D自监督模型的室内、物体级或自动驾驶LiDAR数据差异显著。本文提出Polis,据我们所知首个将草图各向同性高斯正则化(SIGReg)作为原生点云编码器目标的应用,通过涵盖14个城市与建筑尺度数据集的冻结特征基准进行评估。Polis结合几何匹配的余弦不变性、SIGReg以及类似VICReg的抗坍缩项,并采用包含12.8k场景的户外预训练混合数据与保持重力方向的空间视图采样策略。控制消融实验表明,该目标在相同户外数据集上优于师生架构及其他无抗坍缩项的Polis变体。在三个未参与预训练的城市数据集上,Polis在高容量冻结探测下取得23.8%的平均mIoU,优于次优编码器的16.3%;在匹配点数与体素预算下,分别为17.3%与16.1%。对于训练集曾出现在预训练中的数据集,优势依然明显。但在包含细粒度立面与街景标签的局部地面采集数据上,排名反转。结果表明,分布正则化的联合嵌入架构可在挑战性的城市级3D场景中成功应用,且当自监督设计适配该领域采集几何与空间上下文时迁移性能提升,但也揭示了这种专化带来的局限。
原文摘要 · Abstract (English)
Reliable semantic representations derived from city-scale 3D models are increasingly important for urban analysis, infrastructure monitoring, autonomous systems, and heritage conservation. However, urban scenes of large spatial extent captured through aerial surveying differ substantially from the indoor, object-level, and self-driving LiDAR data used to pretrain most 3D self-supervised models. We introduce Polis, to our knowledge the first application of Sketched Isotropic Gaussian Regularization (SIGReg) as an objective for a native point cloud encoder, and evaluate it through a frozen-feature benchmark spanning fourteen city- and building-scale corpora. Polis combines geometrically matched cosine invariance, SIGReg, and VICReg-style anti-collapse terms with a 12.8k-scene outdoor pretraining mixture and gravity-preserving spatial view sampling. Controlled ablations show that this objective outperforms student--teacher architecture alternatives, as well as Polis versions without anti-collapse terms, on the same representative outdoor corpus. On three pretraining-disjoint city datasets, Polis reaches $23.8\%$ mean mIoU versus $16.3\%$ for the next-best encoder under high-capacity frozen probing, and $17.3\%$ versus $16.1\%$ at a matched point and voxel budget. The same city-scale lead holds on datasets whose training sets were seen in pretraining. On localized terrestrial captures with fine-grained facade and streetscape labels, the ranking reverses. Our results show that distributionally-regularized joint embedding architectures can be successful on challenging city-scale 3D scenes, and that transfer improves when self-supervision is designed for the capture geometry and spatial context of this domain while also revealing the limits of this specialization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。