arXiv:2512.14727cs.LGcs.AI2025-12被引 2

小样本下共形预测的统计保证未必实用,医疗场景需警惕

A Critical Perspective on Finite Sample Conformal Prediction Theory in Medical Applications

  • 用校准集生成带统计保证的预测集,但小样本时效果差
  • 实验证明:校准集越小,预测集越大,实用性越低
  • 适合关注小样本医疗模型可靠性的研究者阅读

机器学习正改变医疗,但安全决策需可靠的不确定性估计,而标准模型无法提供。共形预测(CP)可将经验不确定性转化为具有统计保证的估计,通过将模型预测与校准样本结合,生成以任意期望概率包含真实标签的预测集。尽管理论表明校准集大小任意均可保证统计可靠性,但实际应用中,小样本校准集会导致预测集过大,降低可用性。这一问题在数据稀缺的医疗领域尤为关键。我们在医学图像分类任务中进行了实证验证,结果表明:校准集规模直接影响预测集的紧凑性与临床实用性。

原文摘要 · Abstract (English)

Machine learning (ML) is transforming healthcare, but safe clinical decisions demand reliable uncertainty estimates that standard ML models fail to provide. Conformal prediction (CP) is a popular tool that allows users to turn heuristic uncertainty estimates into uncertainty estimates with statistical guarantees. CP works by converting predictions of a ML model, together with a calibration sample, into prediction sets that are guaranteed to contain the true label with any desired probability. An often cited advantage is that CP theory holds for calibration samples of arbitrary size, suggesting that uncertainty estimates with practically meaningful statistical guarantees can be achieved even if only small calibration sets are available. We question this promise by showing that, although the statistical guarantees hold for calibration sets of arbitrary size, the practical utility of these guarantees does highly depend on the size of the calibration set. This observation is relevant in medical domains because data is often scarce and obtaining large calibration sets is therefore infeasible. We corroborate our critique in an empirical demonstration on a medical image classification task.

共形预测医疗AI小样本不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。