arXiv:2605.10029cs.CV2026-05被引 1

用全球统一地表嵌入评估12城贫民窟识别与密度图,发现时间跨年训练最优。

Slum Detection and Density Mapping with AlphaEarth Foundations: A Representation Learning Evaluation Across 12 Global Cities

  • 基于AlphaEarth基础嵌入,用伪标签训练贫民窟检测与密度估计
  • 跨年训练效果优于跨城迁移,10米分辨率难捕捉像素内密度梯度
  • 兴趣点特征提升最大,六城可生成连续多年贫民窟分布图

像素级贫民窟制图长期受限于跨城市泛化能力差、缺乏连续密度估计及全球可比性弱。AlphaEarth Foundations(AEF)是64维、10米分辨率的全球一致年度地表嵌入,为轻量级贫民窟监测提供新基础,但其在贫民窟检测——这一受建筑形态与社会经济过程共同影响的间接任务——中的适用性尚未验证。本研究在12个城市的69个城-年对(2017–2024)上评估AEF在贫民窟分类与亚像素密度估计的表现,采用GRAM伪掩码作为监督标签。涵盖四种训练策略、两种协议(随机划分与3×3空间块交叉验证)、六种辅助特征配置及五种基线模型,并辅以表示层分析(PCA、SHAP)与全区域覆盖推断。五大发现:(1)同城市跨年训练在两种协议下均最优(中位空间F1=0.616,R²=0.466);时间扩展优于跨城迁移,表明城市尺度表征漂移显著;(2)回归R²主要由零/非零边界判别驱动:所有城市正像素R²均为负值,揭示10米分辨率下难以建模像素内密度梯度;(3)PC36在各任务中始终领先;分类在k=32饱和,回归在k=64仍未饱和;(4)兴趣点(POI)特征带来最大密度提升(ΔR²=+0.064);(5)六座满足双任务可用性阈值的城市,全区域覆盖推断(2017–2024)保持贫民窟簇结构稳定(平均SSIM=0.926)。研究厘清了基础模型嵌入在贫民窟监测中的能力边界与互补需求。

原文摘要 · Abstract (English)

Pixel-level slum mapping has long been constrained by limited cross-city generalisation, the absence of continuous density estimation, and weak global comparability. AlphaEarth Foundations (AEF), a globally consistent 64-dimensional annual surface embedding at 10 m, offers a new analysis-ready basis for lightweight slum monitoring, but its applicability to slum detection - an indirectly coupled task shaped by both built form and socio-economic processes - remains untested. We evaluate AEF on slum classification and sub-pixel density estimation across 12 cities and 69 city-year pairs (2017-2024), using GRAM pseudo-masks as supervisory labels. The evaluation spans four training strategies, two protocols (random split and 3x3 spatial block cross-validation), six auxiliary feature configurations, and five baseline models, complemented by representation-level analyses (PCA, SHAP) and full-AOI mapping. Five findings emerge. (1) Same-city cross-year training is optimal under both protocols (median spatial F1 = 0.616, R^2 = 0.466); temporal expansion outperforms cross-city transfer, indicating city-scale representational drift. (2) Regression R^2 is driven primarily by zero/non-zero boundary discrimination: positive-pixel R^2 is consistently negative across all cities, revealing limited capacity to model intra-pixel density gradients at 10 m. (3) PC36 is consistently top-ranked across tasks; classification saturates at k = 32 while regression remains unsaturated at k = 64. (4) POI features yield the largest density gain (Delta R^2 = +0.064). (5) For six cities meeting dual-task usability thresholds, full-AOI inference across 2017-2024 preserves slum cluster structure (mean SSIM = 0.926). The study delineates the capabilities and complementarity needs of foundation-model embeddings for slum monitoring.

贫民窟识别遥感分析基础模型密度映射

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。