用多个模型动态选优,降低医疗AI监控的标注成本。
Active Multiple-Prediction-Powered Inference

- 根据实例特性选择最合适的模型组合,动态分配预测任务。
- 在相同预算下,置信区间比单模型方法窄10%至40%。
- 适合医疗等高成本标注场景,提升模型监控效率。
部署后医疗AI监控需要统计有效且标签高效的评估方法,但由临床医生审阅获取的金标准标签成本高昂。预测驱动推断(PPI)和主动统计推断(ASI)通过结合少量标注样本与大量模型预测,降低了标签成本,但二者均受限于单一预测器,难以适应现代临床流程中存在多个不同成本与准确率的预测器的情况。本文提出主动多预测驱动推断(AM-PPI),将每个实例路由至成本合适的预测器子集,在选定子集的残差不确定性比例下采样金标准标签,并通过重加权预测以最小化估计器方差,所有操作均在单一部署预算下完成。AM-PPI将ASI推广至多预测器场景,并将多预测器PPI从全局分配升级为实例级自适应路由。我们推导出三项决策的闭式Karush-Kuhn-Tucker(KKT)条件,通过双凸性和强对偶性证明,尽管联合问题非联合凸,其固定点仍为全局最优。理论保证渐近正态性、线性预测增强逆倾向加权(AIPW)类中的无偏最小方差性,以及多预测器带来增益的闭式判别准则。在合成数据及三个医疗监控任务上,当路由策略起作用时,AM-PPI的置信区间比单预测器ASI窄10%至40%,其余情况下性能匹配更优基线。
原文摘要 · Abstract (English)
Post-deployment monitoring of healthcare AI requires statistically valid, label-efficient methods, but gold-standard labels from clinician chart review are expensive. Prediction-powered inference (PPI) and active statistical inference (ASI) reduce label cost by combining a small labeled sample with abundant model predictions, but both are restricted to a single predictor, a poor fit for modern clinical pipelines that have multiple predictors of differing cost and accuracy available at inference time. We propose Active Multiple-Prediction-Powered Inference (AM-PPI), which routes each instance to a cost-appropriate predictor subset, samples gold-standard labels in proportion to the chosen subset's residual uncertainty, and reweights predictions to minimize estimator variance, all under a single deployment-time budget. AM-PPI generalizes ASI to leverage multiple predictors and extends Multiple-PPI from global per-predictor allocation to per-instance adaptive routing. We derive closed-form Karush-Kuhn-Tucker (KKT) conditions for all three decisions and prove, via biconvexity and strong duality, that the resulting fixed point is a global optimum despite the joint problem being non-jointly-convex. We establish asymptotic normality with valid coverage, minimum-variance unbiasedness within the linear-prediction augmented inverse propensity weighted (AIPW) class, and a closed-form criterion identifying when multiple predictors help. On synthetic data and three healthcare monitoring tasks, AM-PPI produces 10 to 40 percent narrower confidence intervals (CIs) than single-predictor ASI in the budget regime where routing matters, and matches the better baseline elsewhere.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。