arXiv:2603.16041stat.MEcs.LG2026-03被引 3

AI预测提升统计检验效率,给出所需样本量的精确计算方法

Power Analysis for Prediction-Powered Inference

  • 基于PPI估计量渐近方差推导出闭式功效公式
  • 发现所需样本量减少幅度与预测值和真实值R²成正比
  • 适用于两样本比较等常见场景,适合生物医学研究者使用

现代研究越来越多地利用机器学习和人工智能(AI/ML)模型预测结果,近年来如预测驱动推断(PPI)等方法已发展出有效的下游统计推断程序。然而,经典的功效与样本量公式难以适应此类预测。本文解决一个简单但实用的问题:在新AI/ML模型具有高预测能力的前提下,为达到预期统计功效,需要多少带标签样本?通过刻画PPI估计量的渐近方差,并应用Wald检验反演,我们推导出闭式功效公式,覆盖两样本比较和2×2表中的风险度量等广泛情形。结果显示,相对于经典设计,所需带标签样本量的减少程度大致与预测值和真实值之间的R²成正比。分析公式经蒙特卡洛模拟验证,并在三个当代生物医学应用中展示:单细胞转录组学、临床血压测量及皮肤镜影像。相关软件以R包和在线计算器形式开源,地址见https://github.com/yiqunchen/pppower。

原文摘要 · Abstract (English)

Modern studies increasingly leverage outcomes predicted by machine learning and artificial intelligence (AI/ML) models, and recent work, such as prediction-powered inference (PPI), has developed valid downstream statistical inference procedures. However, classical power and sample size formulas do not readily account for these predictions. In this work, we tackle a simple yet practical question: given a new AI/ML model with high predictive power, how many labeled samples are needed to achieve a desired level of statistical power? We derive closed-form power formulas by characterizing the asymptotic variance of the PPI estimator and applying Wald test inversion to obtain the required labeled sample size. Our results cover widely used settings including two-sample comparisons and risk measures in 2x2 tables. We find that a useful rule of thumb is that the reduction in required labeled samples relative to classical designs scales roughly with the R2 between the predictions and the ground truth. Our analytical formulas are validated using Monte Carlo simulations, and we illustrate the framework in three contemporary biomedical applications spanning single-cell transcriptomics, clinical blood pressure measurement, and dermoscopy imaging. We provide our software as an R package and online calculators at https://github.com/yiqunchen/pppower.

统计推断AI预测功效分析生物医学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。