arXiv:2502.14166stat.MLcs.LG2025-02ICML被引 8

用机器学习预测结果自适应调整多均值估计,提升大规模统计推断精度。

Prediction-Powered Adaptive Shrinkage Estimation

  • 利用预测偏差校正与跨任务信息共享,实现多均值的自适应收缩估计
  • 通过最小化无偏风险估计自动确定收缩程度,在真实和合成数据上均更优
  • 适合需要同时处理大量统计问题且有可靠预测模型的场景

Prediction-Powered Inference (PPI) 是一种通过结合少量高质量数据与机器学习(ML)预测来增强统计估计的强大框架。尽管以往研究已证明 PPI 在单一统计问题中的优势,但现代应用需解决大量并行的统计问题。本文提出 Prediction-Powered Adaptive Shrinkage (PAS),将 PPI 与经验贝叶斯收缩方法结合,用于改进多个均值的估计。PAS 在每个任务中对噪声较大的 ML 预测进行去偏,并利用这些预测作为收缩参考点,跨任务借力。收缩量由无偏风险估计的最小化决定,我们证明该调参策略渐近最优。在合成与真实数据集上的实验表明,PAS 能自适应 ML 预测的可靠性,在大规模应用中优于传统及现代基线方法。

原文摘要 · Abstract (English)

Prediction-Powered Inference (PPI) is a powerful framework for enhancing statistical estimates by combining limited gold-standard data with machine learning (ML) predictions. While prior work has demonstrated PPI's benefits for individual statistical problems, modern applications require answering numerous parallel statistical questions. We introduce Prediction-Powered Adaptive Shrinkage (PAS), a method that bridges PPI with empirical Bayes shrinkage to improve the estimation of multiple means. PAS debiases noisy ML predictions within each task and then borrows strength across tasks by using those same predictions as a reference point for shrinkage. The amount of shrinkage is determined by minimizing an unbiased estimate of risk, and we prove that this tuning strategy is asymptotically optimal. Experiments on both synthetic and real-world datasets show that PAS adapts to the reliability of the ML predictions and outperforms traditional and modern baselines in large-scale applications.

统计推断机器学习收缩估计多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。