对比了12种模型在结直肠病理图像分类中的表现,发现基于Transformer的模型效果最好。
Performance and Interpretability of Convolutional, Transformer, and Hybrid Deep Learning Models in Colorectal Histology Classification
- 采用标准微调流程,对比12种预训练模型在结直肠病理图像上的分类性能。
- 最佳模型准确率达97.1%,所有模型均超过93%,复杂间质类最难识别。
- 适合关注病理图像自动分析的医生、研究人员及医学AI开发者参考。
深度学习已成为计算病理学的重要工具,实现对组织病理图像的自动化分析。尽管卷积神经网络(CNN)长期主导该领域,但基于Transformer和混合架构的模型近年表现出良好潜力。然而,针对结直肠病理学的全面比较仍有限。本研究评估了十二种ImageNet预训练的CNN、Transformer及混合架构,使用包含5000张图像切片的Kather结直肠病理数据集,涵盖八种类别。所有模型均采用标准化的迁移学习与微调协议,并通过准确率、精确率、敏感度、特异度、F1分数、ROC-AUC、Cohen's kappa和马修斯相关系数等多指标评估。所有模型均表现优异,准确率介于93.2%至97.1%之间。EVA-02达到最高综合性能(准确率97.1%,F1分数97.0%),紧随其后的是ViT-B/16。在CNN中,ResNet34和ConvNeXt-Tiny表现突出,准确率分别为96.4%和96.3%。总体而言,Transformer架构在各项指标上表现最强,但与最优CNN模型的差距较小。每类分析显示所有组织类别均有稳健分类能力,其中复杂间质类最具挑战性。结果表明,基于Transformer的架构具有最高预测性能,而现代CNN则在精度与模型复杂度间取得良好平衡。该研究为结直肠病理图像分类的主要深度学习范式提供了全面基准。
原文摘要 · Abstract (English)
Deep learning has become an important tool in computational pathology, enabling automated analysis of histopathological images. While convolutional neural networks (CNNs) have traditionally dominated this field, transformer-based and hybrid architectures have recently demonstrated promising performance. However, comprehensive comparisons of these approaches for colorectal histopathology remain limited. This study evaluated twelve ImageNet-pretrained CNN, transformer, and hybrid architectures using the Kather colorectal histopathology dataset containing 5,000 image tiles from eight tissue classes. All models were trained using a standardized transfer-learning and fine-tuning protocol and assessed using multiple performance metrics, including accuracy, precision, sensitivity, specificity, F1-score, ROC-AUC, Cohen's kappa, and Matthews correlation coefficient. All evaluated models achieved high classification performance, with accuracies ranging from 93.2% to 97.1%. EVA-02 achieved the highest overall performance (97.1% accuracy, 97.0% F1-score), closely followed by ViT-B/16. Among CNNs, ResNet34 and ConvNeXt-Tiny demonstrated highly competitive performance, achieving accuracies of 96.4% and 96.3%, respectively. Transformer architectures generally produced the strongest results across evaluation metrics, although the performance gap between the best transformer and CNN models was relatively small. Per-class analysis showed consistently strong classification performance across all tissue categories, with Complex Stroma representing the most challenging class. Overall, transformer-based architectures achieved the highest predictive performance, whereas modern CNNs provided a favorable balance between accuracy and model complexity. These findings provide a comprehensive benchmark of major deep learning paradigms for colorectal histopathology classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。