用机器学习预测补全数据后,仍能给出可靠置信区间。
Prediction-Powered Inference with Imputed Covariates and Nonuniform Sampling
- 通过自举法构建置信区间,适配非均匀抽样和部分特征缺失场景。
- 无需假设模型质量,结果比不使用预测的方法更紧致。
- 适用于卫星图像、语言模型等下游统计分析,无需额外计算。
机器学习模型生成的预测正被广泛用于后续统计分析中,如基于卫星图像的经济与环境指标预测、语言模型对社会科学研究中人类评分的近似。然而,若未正确处理预测误差,传统统计方法将失效。已有研究提出‘预测-再校正’估计器,在小规模完整样本下可提供有效置信区间。本文拓展该方法,引入适用于非均匀样本(加权、分层或聚类)及任意子集特征缺失情形的自举置信区间。该方法在不依赖模型性能假设的前提下保证有效性,且置信区间宽度不劣于不使用机器学习预测的方法。
原文摘要 · Abstract (English)
Machine learning models are increasingly used to produce predictions that serve as input data in subsequent statistical analyses. For example, computer vision predictions of economic and environmental indicators based on satellite imagery are used in downstream regressions; similarly, language models are widely used to approximate human ratings and opinions in social science research. However, failure to properly account for errors in the machine learning predictions renders standard statistical procedures invalid. Prior work uses what we call the Predict-Then-Debias estimator to give valid confidence intervals when machine learning algorithms impute missing variables, assuming a small complete sample from the population of interest. We expand the scope by introducing bootstrap confidence intervals that apply when the complete data is a nonuniform (i.e., weighted, stratified, or clustered) sample and to settings where an arbitrary subset of features is imputed. Importantly, the method can be applied to many settings without requiring additional calculations. We prove that these confidence intervals are valid under no assumptions on the quality of the machine learning model and are no wider than the intervals obtained by methods that do not use machine learning predictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。