用地理编码和质量评估提升卫星图像贫困预测精度
Enhancing the KidSat Model: Integrating Geographical Encoding and Data Quality Assessment for Childhood Poverty Prediction

- 通过重聚类降低标签维度,融合球谐函数地理特征
- 图像质量筛选后,贫困预测平均绝对误差降18.83%
- 适合做卫星遥感社会经济预测的科研与政策人员
利用卫星图像进行精准贫困制图常受三方面制约:(i)调查监督数据噪声多、稀疏;(ii)图像质量差如云层遮挡;(iii)纯图像模型缺乏显式空间结构。基于KidSat框架,本文提出增强管道:首先通过人口普查数据重聚类,将标签维度从103降至51;其次引入两阶段质量筛查,剔除严重云遮或损坏样本;最后融合DINOv2视觉嵌入与球谐函数(SH)位置特征。实验显示,该方法使集群层面严重剥夺比例预测的平均绝对误差从0.2167降至0.1759,相对降低18.83%。扩展至33个非洲国家时,最优配置整体MAE达0.1658。结果表明SH特征始终提升性能,而更高容量的坐标网络(SH+SIREN)在无合理目标设计下反而表现更差。梯度提升树头(XGBoost/LightGBM)能最好捕捉融合表示中的非线性关系。本研究为仅使用公开数据的卫星社会经济预测提供了可扩展、有依据的方法。
原文摘要 · Abstract (English)
Accurate poverty mapping using satellite imagery is often hindered by (i) noisy and sparse survey-derived supervision, (ii) image quality issues such as cloud cover and image corruption, and (iii) lack of explicit spatial structure in image-only models. Building on the KidSat framework, we develop an enhanced pipeline that improves predictive accuracy via refined data preprocessing, systematic image quality assessment, and mathematically defined geographic encoding. First, we refine the fine-tuning target matrix by resolving high-cardinality sparsity and reducing one-hot dimensionality from 103 to 51 via DHS re-aggregation. Second, we introduce a simple two-stage quality-screening procedure to filter heavily clouded or corrupted observations. Third, we fuse DINOv2 visual embeddings with Spherical Harmonics (SH) location features. Across extensive experiments, these changes reduce MAE from 0.2167 to 0.1759, corresponding to an 18.83% relative reduction on the cluster-level severe-deprivation proportion scale. When extended from 16 to 33 African countries, the best-performing configuration achieves an overall MAE of 0.1658. We find that SH features consistently improve performance over the image-only backbone, whereas higher-capacity coordinate Multi Layer Perception augmentation (SH+SIREN) can underperform without carefully designed objectives. Finally, gradient-boosted tree heads (XGBoost/LightGBM) most effectively exploit nonlinear interactions in the fused visual-geographic representation. These findings provide a scalable and principled recipe for improving satellite-based socioeconomic predictions using only publicly accessible data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。