用机器学习预测提升置信区间的效率,且全程有效。
Anytime-valid, Bayes-assisted, Prediction-Powered Inference
- 结合预测与贝叶斯先验,构建随时间递增的置信序列。
- 在真实与合成数据上验证,比传统方法更高效且始终可靠。
- 适合需要持续更新、高时效性的统计推断场景。
在大量未标注数据和少量标注数据的前提下,预测驱动推断(PPI)利用机器学习预测结果,提升仅基于标注数据的置信区间估计效率,同时保持固定时间有效性。本文将PPI框架扩展至序贯设置,即标注与未标注数据随时间增长。借助Ville不等式与混合方法,提出预测驱动置信序列,该方法在时间上一致渐近有效,并自然融入对预测质量的先验知识以进一步提升效率。文中详述了方法设计思路,并在真实与合成数据上展示了其有效性。
原文摘要 · Abstract (English)
Given a large pool of unlabelled data and a smaller amount of labels, prediction-powered inference (PPI) leverages machine learning predictions to increase the statistical efficiency of confidence interval procedures based solely on labelled data, while preserving fixed-time validity. In this paper, we extend the PPI framework to the sequential setting, where labelled and unlabelled datasets grow over time. Exploiting Ville's inequality and the method of mixtures, we propose prediction-powered confidence sequence procedures that are asymptotically valid uniformly over time and naturally accommodate prior knowledge on the quality of the predictions to further boost efficiency. We carefully illustrate the design choices behind our method and demonstrate its effectiveness in real and synthetic examples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。