对比基础模型与专用模型在前列腺癌诊断中的表现,发现数据充足时专用模型更优。
Foundation Models -- A Panacea for Artificial Intelligence in Pathology?
- 用大规模数据对比基础模型与端到端专用模型的性能差异。
- 数据足够时,专用模型准确率更高,误诊率显著降低。
- 适合临床部署的病理AI需重视任务定制化训练,而非依赖通用模型。
人工智能在病理学中的角色已从辅助诊断发展为揭示全切片图像(WSIs)中的预测性形态模式。近期,基于自监督预训练的基础模型(FMs)被视为解决多样化下游任务的通用方案。然而,其临床适用性与泛化优势仍存疑问。本文聚焦于前列腺癌诊断与格里森分级的临床级人工智能,使用来自11个国家15个机构、7,342名患者超过10万例核心针活检样本进行最大规模验证。在多实例学习框架下,比较两种基础模型与一个完全端到端的任务专用(TS)模型。结果挑战了基础模型普遍优于专用模型的假设:在数据稀缺时,基础模型表现良好;但当标注数据充足时,其性能趋于收敛,甚至被专用模型超越。值得注意的是,充分的任务专用训练显著降低了临床显著误判、复杂形态误诊及不同扫描仪间的差异。此外,基础模型能耗最高达专用模型的35倍,引发可持续性担忧。研究强调,尽管基础模型利于快速原型设计,但其作为临床可用医学AI通用解法的地位尚不确定。对于高风险临床应用,严格验证与任务定制化训练至关重要。建议融合基础模型与端到端学习优势,构建鲁棒且资源高效的临床级病理AI解决方案。
原文摘要 · Abstract (English)
The role of artificial intelligence (AI) in pathology has evolved from aiding diagnostics to uncovering predictive morphological patterns in whole slide images (WSIs). Recently, foundation models (FMs) leveraging self-supervised pre-training have been widely advocated as a universal solution for diverse downstream tasks. However, open questions remain about their clinical applicability and generalization advantages over end-to-end learning using task-specific (TS) models. Here, we focused on AI with clinical-grade performance for prostate cancer diagnosis and Gleason grading. We present the largest validation of AI for this task, using over 100,000 core needle biopsies from 7,342 patients across 15 sites in 11 countries. We compared two FMs with a fully end-to-end TS model in a multiple instance learning framework. Our findings challenge assumptions that FMs universally outperform TS models. While FMs demonstrated utility in data-scarce scenarios, their performance converged with - and was in some cases surpassed by - TS models when sufficient labeled training data were available. Notably, extensive task-specific training markedly reduced clinically significant misgrading, misdiagnosis of challenging morphologies, and variability across different WSI scanners. Additionally, FMs used up to 35 times more energy than the TS model, raising concerns about their sustainability. Our results underscore that while FMs offer clear advantages for rapid prototyping and research, their role as a universal solution for clinically applicable medical AI remains uncertain. For high-stakes clinical applications, rigorous validation and consideration of task-specific training remain critically important. We advocate for integrating the strengths of FMs and end-to-end learning to achieve robust and resource-efficient AI pathology solutions fit for clinical use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。