arXiv:2601.20819stat.MLcs.LG2026-01被引 5

用预测数据提升统计效率,同时纠正偏差,让模型预测更可信。

Demystifying Prediction Powered Inference

  • 将大样本预测结果与小样本标签结合,通过偏差校正提升推断精度。
  • 实验证明PPI比传统方法区间更窄,但重复使用数据会导致覆盖不足。
  • 提供诊断工具和决策流程图,帮助研究者选对方法避免误用。

机器学习预测在生物医学、环境科学和社会科学中被广泛用于补充不完整或成本高昂的观测数据。然而,将预测当作真实值会引入偏差,忽略预测则浪费信息。预测驱动推断(PPI)提供了一个系统框架,利用大规模无标签数据的预测结果来提高统计效率,并通过小规模有标签子集显式校正偏差,确保推断有效性。尽管PPI变体不断增多,其细微差异使实践者难以判断何时何地合理应用。本文通过整合理论基础、方法扩展、与经典统计学的联系以及诊断工具,构建统一实用的工作流程。基于MOSAIKS住房价格数据,我们发现不同PPI方法可产生更紧凑的置信区间,但重复使用训练数据(双采样)会导致置信区间过窄、覆盖概率低于名义水平。在非随机缺失机制下,所有方法(包括仅用标签数据的经典推断)均产生偏估计。本文提供决策流程图,将假设违背与适用的PPI变体对应;汇总代表性方法表;并提出评估核心假设的实用诊断策略。将PPI视为通用范式而非单一估计器,本工作弥合了方法创新与实际应用之间的鸿沟,助力研究者负责任地融合预测与有效推断。

原文摘要 · Abstract (English)

Machine learning predictions are increasingly used to supplement incomplete or costly-to-measure outcomes in fields such as biomedical research, environmental science, and social science. However, treating predictions as ground truth introduces bias while ignoring them wastes valuable information. Prediction-Powered Inference (PPI) offers a principled framework that leverages predictions from large unlabeled datasets to improve statistical efficiency while maintaining valid inference through explicit bias correction using a smaller labeled subset. Despite its potential, the growing PPI variants and the subtle distinctions between them have made it challenging for practitioners to determine when and how to apply these methods responsibly. This paper demystifies PPI by synthesizing its theoretical foundations, methodological extensions, connections to existing statistics literature, and diagnostic tools into a unified practical workflow. Using the MOSAIKS housing price data, we show that PPI variants produce tighter confidence intervals than complete-case analysis, but that double-dipping, i.e. reusing training data for inference, leads to anti-conservative confidence intervals and below-nominal coverage. Under missing-not-at-random mechanisms, all methods, including classical inference using only labeled data, yield biased estimates. We provide a decision flowchart linking assumption violations to appropriate PPI variants, a summary table of representative methods, and practical diagnostic strategies for evaluating core assumptions. By framing PPI as a general recipe rather than a single estimator, this work bridges methodological innovation and applied practice, helping researchers responsibly integrate predictions into valid inference.

预测推断统计校正数据缺失方法指南

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。