解决预测标签下空间依赖数据的稳定推断问题
Spatially Robust Inference with Predicted and Missing at Random Labels
- 提出双重稳健估计器结合交叉拟合,处理缺失随机标签
- 发现交叉拟合引入折叠相关性,导致方差估计失真
- 设计空间异方差自相关一致校正,提升小样本校准效果
当结果数据获取成本高时,研究者常以机器学习模型预测替代未标注样本,这对后续统计推断带来影响。现有方法在独立采样下可提供有效不确定性量化,但真实场景常存在缺失随机(MAR)标签与空间依赖。本文提出一种带交叉拟合扰动项的双重稳健估计器。研究表明,交叉拟合会引入折叠层级相关性,扭曲空间方差估计,导致置信区间不稳定或过于保守。为此,我们提出一种刀切空间异方差自相关一致(jackknife spatial HAC)方差校正方法,将空间依赖与折叠诱导噪声分离。在标准识别与依赖条件下,所得区间渐近有效。模拟与基准数据集显示,该方法在有限样本下校准性能显著提升,尤其在MAR标签与聚类采样情形下表现突出。
原文摘要 · Abstract (English)
When outcome data are expensive or onerous to collect, scientists increasingly substitute predictions from machine learning and AI models for unlabeled cases, a process which has consequences for downstream statistical inference. While recent methods provide valid uncertainty quantification under independent sampling, real-world applications involve missing at random (MAR) labeling and spatial dependence. For inference in this setting, we propose a doubly robust estimator with cross-fit nuisances. We show that cross-fitting induces fold-level correlation that distorts spatial variance estimators, producing unstable or overly conservative confidence intervals. To address this, we propose a jackknife spatial heteroscedasticity and autocorrelation consistent (HAC) variance correction that separates spatial dependence from fold-induced noise. Under standard identification and dependence conditions, the resulting intervals are asymptotically valid. Simulations and benchmark datasets show substantial improvement in finite-sample calibration, particularly under MAR labeling and clustered sampling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。