提出三项可靠性指标,验证病理图像分类中MIL模型的可信度
Beyond accuracy: quantifying the reliability of Multiple Instance Learning for Whole Slide Image classification
- 设计三类量化指标评估MIL模型在病理图像中的可靠性
- 发现均值池化实例模型可靠性最优,且结构简单高效
- 适合关注模型可信度的临床医疗AI研究者参考
机器学习模型虽已广泛应用于多个领域,但其可靠性——即生成可信赖、一致且符合领域规律的预测结果——仍是关键挑战。针对计算病理学中用于全切片图像(WSI)分类的多重实例学习(MIL)模型,现有研究极少从可靠性角度进行评估,导致其在临床决策等高风险场景中的适用性存疑。本文提出三项定量可靠性评估指标,并在三个基于区域标注的病理数据集上对多种主流MIL架构进行了测试。结果显示,尽管结构简单且计算高效,均值池化实例(MEAN-POOL-INS)模型在可靠性方面表现最优。该研究强调在评估MIL模型时需兼顾预测性能与可靠性,确立了MEAN-POOL-INS作为未来研究的可靠基线。
原文摘要 · Abstract (English)
Machine learning models have become integral to many fields, but their reliability, defined as producing dependable, trustworthy, and domain-consistent predictions, remains a critical concern. Multiple Instance Learning (MIL) models designed for Whole Slide Image (WSI) classification in computational pathology are rarely evaluated in terms of reliability, leaving a key gap in understanding their suitability for high-stakes applications like clinical decision-making. In this paper, we address this gap by introducing three quantitative metrics for reliability assessment and applying them to several widely used MIL architectures across three region-wise annotated pathology datasets. Our findings indicate that the mean pooling instance (MEAN-POOL-INS)model demonstrates superior reliability compared to other networks, despite its simple architectural design and computational efficiency. These findings underscore the need of reliability evaluation alongside predictive performance in MIL models and establish MEAN-POOL-INS as a strong, trustworthy baseline for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。