用逆概率加权改进预测推断,让小样本标签也能准确估算整体数据。
Prediction-Powered Inference with Inverse Probability Weighting
- 通过逆概率加权修正标签选择偏差,结合模型预测与真实标签。
- 估计标签概率后,仍能保持95%置信度和预测带来的方差降低优势。
- 适合做带标签数据稀缺的统计推断,尤其在标签非随机时有用。
预测推断(PPI)是一种针对部分标注数据的有效统计推断框架,将大规模未标注数据上的模型预测与小规模标注子集的偏差校正相结合。本文基于协变量偏移下的现有PPI结果,表明PPI校正可直接从设计基础视角理解,并可通过霍夫里茨-汤普森与哈杰克型校正自然处理信息性标注。这一联系将设计基础抽样思想与现代预测辅助推断统一起来,使得当单位间标注概率不同时,估计量依然有效。我们考虑常见情形:包含概率未知,但由正确设定的模型估计。模拟结果显示,使用估计倾向值的IPW调整PPI性能接近已知概率情况,既保持名义覆盖率,也保留了PPI的方差缩减优势。
原文摘要 · Abstract (English)
Prediction-powered inference (PPI) is a recent framework for valid statistical inference with partially labeled data, combining model-based predictions on a large unlabeled set with bias correction from a smaller labeled subset. Building on existing PPI results under covariate shift, we show that PPI rectification admits a direct design-based interpretation, and that informative labeling can be handled naturally by Horvitz--Thompson and Hájek-style corrections. This connection unites design-based survey sampling ideas with modern prediction-assisted inference, yielding estimators that remain valid when labeling probabilities vary across units. We consider the common setting where the inclusion probabilities are not known but estimated from a correctly specified model. In simulations, the performance of IPW-adjusted PPI with estimated propensities closely matches the known-probability case, retaining both nominal coverage and the variance-reduction benefits of PPI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。