用16项任务训练出高效病理基础模型,仅需6%数据即达自监督水平
Tissue Concepts: supervised foundation models in computational pathology
- 通过16个任务联合训练,构建可复用的组织特征编码器
- 仅用91.2万张切片的6%数据,性能媲美自监督模型
- 适用于跨中心病理图像分析,尤其适合数据有限的场景
由于病理科医生工作负荷增加,自动化辅助诊断和定量生物标志物评估的需求日益迫切。基础模型有望提升模型在不同中心间的泛化能力,并作为数据高效开发专用且稳健AI模型的起点。然而,训练基础模型通常需要大量数据、计算资源和时间。本文提出一种监督训练方法,显著降低这些成本。该方法基于多任务学习,利用总计91.2万张切片上的16种分类、分割与检测任务联合训练一个共享编码器,我们称其为组织概念编码器(Tissue Concepts encoder)。为评估其跨中心性能与泛化能力,使用四种最常见实体癌——乳腺、结肠、肺和前列腺——的全切片图像进行分类实验。结果表明,组织概念模型在性能上可与自监督训练模型相媲美,但仅需其6%的训练切片数量。此外,该编码器在域内和域外数据上均优于ImageNet预训练编码器。
原文摘要 · Abstract (English)
Due to the increasing workload of pathologists, the need for automation to support diagnostic tasks and quantitative biomarker evaluation is becoming more and more apparent. Foundation models have the potential to improve generalizability within and across centers and serve as starting points for data efficient development of specialized yet robust AI models. However, the training foundation models themselves is usually very expensive in terms of data, computation, and time. This paper proposes a supervised training method that drastically reduces these expenses. The proposed method is based on multi-task learning to train a joint encoder, by combining 16 different classification, segmentation, and detection tasks on a total of 912,000 patches. Since the encoder is capable of capturing the properties of the samples, we term it the Tissue Concepts encoder. To evaluate the performance and generalizability of the Tissue Concepts encoder across centers, classification of whole slide images from four of the most prevalent solid cancers - breast, colon, lung, and prostate - was used. The experiments show that the Tissue Concepts model achieve comparable performance to models trained with self-supervision, while requiring only 6% of the amount of training patches. Furthermore, the Tissue Concepts encoder outperforms an ImageNet pre-trained encoder on both in-domain and out-of-domain data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。