医学视觉大模型的不确定性评估,关键在选对预训练数据和方法。
Uncertainty of Vision Medical Foundation Models

- 用领域特定数据预训练+自监督学习,点预测更可靠。
- 不同数据训练的模型,校准后仍存在不确定性差异。
- 领域专用模型可提升置信区域预测效率,适合临床部署。
在高风险医疗场景中,准确的不确定性估计至关重要。传统方法依赖模型输出的概率(点预测),缺乏预测覆盖率的严格保证,常需额外校准。相比之下,合取预测(区域预测)提供有限样本下的有效性保障,确保真实值以指定置信水平包含于预测集合中。本研究系统评估了预训练方式、数据规模与领域对点预测与区域预测不确定性量化的影响,比较了眼科、病理学及胸部X光领域的专用视觉医学基础模型与通用视觉基础模型。实验覆盖多个在视网膜、组织病理学和胸部X光数据上训练的基础模型,并应用多种校准技术。结果表明:(1) 在高质量领域数据上结合自监督学习进行预训练,可获得比通用领域预训练更优的点预测校准效果;(2) 标准重校准方法无法完全消除不同数据源训练模型间的不确定性差异;(3) 领域专用基础模型能实现更高效的合取预测。研究强调应综合考虑点预测与区域预测,推动医疗AI系统可靠性与可信度提升。
原文摘要 · Abstract (English)
Accurate uncertainty estimation is essential for machine learning systems de- ployed in high-stakes domains such as medicine. Traditional approaches primarily rely on probability outputs from trained models (point predictions), which provide no formal guarantees on prediction coverage and often require additional calibra- tion techniques to improve reliability. In contrast, conformal prediction (region prediction) offers a principled alternative by generating prediction sets with finite- sample validity guarantees, ensuring that the ground truth is contained within the set at a specified confidence level. In this study, we explore the impact of pre-training approach, dataset scale and domain on both point and region-level uncertainty quantification, by studying domain-specific vision medical foundation models vs. general domain vision foundation models. We conduct a comprehensive evaluation across foundation models trained on retinal, histopathological, and Chest X-Rays data, applying various calibration techniques. Our results demonstrate that (1) pre-training on higher-quality domain-specific datasets along with self-supervised learning leads to better-calibrated point predictions than general domain pre-training, (2) stan- dard re-calibration methods alone cannot fully mitigate uncertainty discrepancies across models trained on different data sources, (3) domain-specific foundation model can lead to more efficient conformal prediction. These findings highlight the importance of careful model selection and the inte- gration of both point and region prediction to enhance the reliability and trust- worthiness of medical AI systems. Our work underscores the need for a holistic approach to uncertainty quantification in recent development of medical vision foundation model, ensuring robust and interpretable AI-driven decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。