用卫星图像预测贫困后,提出两种无需新数据的校正方法,让一张地图可支持多次政策评估。
Debiasing Machine Learning Predictions for Causal Inference Without Additional Ground Truth Data: "One Map, Many Trials" in Satellite-Driven Poverty Analysis
- 提出线性校正和Tweedie去收缩法,修复模型预测偏差
- 在无额外真实数据条件下,使因果效应估计几乎无偏
- 适用于卫星、污染、人口等重复使用的预测场景
基于地球观测(EO)数据的机器学习模型在预测家庭财富指数方面表现优异,可生成高分辨率财富地图,支持多轮因果试验,缓解全球发展研究中的数据匮乏问题。但标准训练目标追求整体准确率,导致预测值向均值收缩,削弱了因果效应估计,限制了政策评估应用。现有去偏方法如预测驱动推断(PPI)需额外收集真实标签数据,难以在数据稀缺环境中使用。本文提出两种后处理校正方法:线性校正(LCC)在保留集上估算线性变换;Tweedie校正通过密度得分与上游学习的噪声尺度局部去收缩预测。我们提供实用诊断工具判断是否需要校正,并讨论局限性。在分析结果、模拟及人口健康调查(DHS)数据实验中,两种方法均显著降低衰减效应;其中Tweedie校正使处理效应估计近乎无偏,实现“一张地图,多次试验”的范式。尽管以EO-ML财富制图为背景,方法不局限于地理空间,适用于任何重复使用插补结果的场景(如污染指数、人口密度或大模型衍生指标)。
原文摘要 · Abstract (English)
Machine learning models trained on Earth observation data, such as satellite imagery, have demonstrated significant promise in predicting household-level wealth indices, enabling the creation of high-resolution wealth maps that can be leveraged across multiple causal trials while addressing chronic data scarcity in global development research. However, because standard training objectives prioritize overall predictive accuracy, these predictions often suffer from shrinkage toward the mean, leading to attenuated estimates of causal treatment effects and limiting their utility in policy evaluations. Existing debiasing methods, such as Prediction-Powered Inference (PPI), can handle this attenuation bias but require additional fresh ground-truth data at the downstream stage of causal inference, which restricts their applicability in data-scarce environments. We introduce and evaluate two post-hoc correction methods -- Linear Calibration Correction (LCC) and a Tweedie's correction approach -- that substantially reduce shrinkage-induced prediction bias without relying on newly collected labeled data. LCC applies a simple linear transformation estimated on a held-out calibration split; Tweedie's method locally de-shrink predictions using density score estimates and a noise scale learned upstream. We provide practical diagnostics for when a correction is warranted and discuss practical limitations. Across analytical results, simulations, and experiments with Demographic and Health Surveys (DHS) data, both approaches reduce attenuation; Tweedie's correction yields nearly unbiased treatment-effect estimates, enabling a "one map, many trials" paradigm. Although we demonstrate on EO-ML wealth mapping, the methods are not geospatial-specific: they apply to any setting where imputed outcomes are reused downstream (e.g., pollution indices, population density, or LLM-derived indicators).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。