arXiv:2602.04525cs.CVcs.AI2026-02

用半监督方法提升城市贫民窟遥感图像分割精度,兼顾数据质量评估。

SLUM-i: Semi-supervised Learning for Urban Mapping of Informal Settlements and Data Quality Benchmarking

  • 通过视觉基础模型筛选未标注样本,结合类别感知阈值动态生成伪标签。
  • 在10%~30%标注预算下,平均交并比提升5.9个百分点,30%时超监督基线。
  • 构建跨七城多洲基准数据集,支持贫民窟识别与数据质量分析。

超高分辨率遥感影像为划定贫民窟提供了可扩展基础,但稀疏标注、贫民窟与背景像素严重不平衡,以及城市间形态和图像-掩码对应关系的异质性,增加了模型开发难度。本文提出SLUM-i,一个半监督语义分割框架,并建立涵盖三大洲七座城市的地球观测基准。该基准包含新标注的拉合尔数据集及配套的卡拉奇、孟买数据集,外加四个公开城市数据集,共14,458个RGB图像-掩码块。我们通过类别构成、边界形态、灰度可分性、掩码边界与图像边缘对应度及特征分布差异量化跨城市异质性。为实现标签高效映射,SLUM-i采用表示引导的未标注池筛选(基于DINOv2-Small嵌入移除最不相似块),以及类别感知自适应阈值(通过全局均置信度和每类均软最大指数移动平均调整伪标签接受率)。在10%、20%、30%标注预算下,使用卷积网络与视觉变换器骨干,五次随机种子实验表明,在多个城市-预算组合中优于相应UniMatch基线,平均交并比最高提升5.9个百分点。30%预算下,ResNet-101配置在七个城市中的四个达到或超过全标注监督基线。两项组件仅在训练阶段运行,无推理开销。

原文摘要 · Abstract (English)

Very-high-resolution remote-sensing imagery provides a scalable basis for delineating informal settlements, but sparse annotations, severe imbalance between informal-settlement and background pixels, and cross-city heterogeneity in urban morphology and image--mask correspondence complicate model development. We present SLUM-i, a semi-supervised semantic segmentation framework together with a geographically diverse seven-city Earth observation benchmark spanning three continents. The benchmark combines a newly annotated Lahore dataset and companion Karachi and Mumbai datasets with four publicly released city datasets, totaling 14,458 RGB image--mask tiles. We quantify cross-city heterogeneity using class composition, boundary morphology, grayscale separability, correspondence between mask boundaries and image edges, and divergence between measured feature distributions. For label-efficient mapping, SLUM-i combines representation-guided unlabeled-pool curation, using embeddings from a vision foundation model (DINOv2-Small) to remove the least-similar tiles, and Class-Aware Adaptive Thresholding, which adapts pseudo-label acceptance by class through global mean-confidence and per-class mean-softmax exponential moving averages. Experiments at 10%, 20%, and 30% label budgets, using convolutional and vision-transformer backbones and five random seeds, demonstrate improvements over the corresponding UniMatch baselines in multiple city--budget settings, reaching +5.9 percentage points in mean intersection-over-union. At the 30% budget, the ResNet-101 configuration matches or exceeds its corresponding fully labeled supervised baseline in four of seven cities. Both components operate only during training and add no inference overhead.

遥感分割半监督学习贫民窟识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。