对比多种深度模型在膀胱癌图像分类中的表现与可解释性。
Deep Modeling and Interpretation for Bladder Cancer Classification
- 用13种模型比较卷积与视觉变换器在膀胱癌图像上的分类能力。
- ConvNext准确率约60%,ViT类模型校准效果更好。
- ViT适合解释分布外样本,适用于临床诊断场景。
基于视觉变换器(ViT)和卷积神经网络(CNN)的深度模型在自然图像数据集上表现优异,但在医学影像中,病灶区域通常占图像比例极小,导致性能受限。为此,本研究评估了最新深度模型在膀胱癌分类任务中的表现。采用13种模型(4种CNN、8种Transformer)进行标准分类,分析模型校准性,并使用GradCAM++评估可解释性。在公开多中心膀胱癌数据集上模拟约300次实验,结果表明:ConvNext系列在分类中泛化能力有限,准确率约为60%;而ViT类模型校准效果优于ConvNext和Swin Transformer系列。通过测试时增强进一步提升可解释性。最终发现,无模型能兼顾所有场景:ConvNext适合分布内样本,而ViT及其变体更适用于解释分布外样本。
原文摘要 · Abstract (English)
Deep models based on vision transformer (ViT) and convolutional neural network (CNN) have demonstrated remarkable performance on natural datasets. However, these models may not be similar in medical imaging, where abnormal regions cover only a small portion of the image. This challenge motivates this study to investigate the latest deep models for bladder cancer classification tasks. We propose the following to evaluate these deep models: 1) standard classification using 13 models (four CNNs and eight transormer-based models), 2) calibration analysis to examine if these models are well calibrated for bladder cancer classification, and 3) we use GradCAM++ to evaluate the interpretability of these models for clinical diagnosis. We simulate $\sim 300$ experiments on a publicly multicenter bladder cancer dataset, and the experimental results demonstrate that the ConvNext series indicate limited generalization ability to classify bladder cancer images (e.g., $\sim 60\%$ accuracy). In addition, ViTs show better calibration effects compared to ConvNext and swin transformer series. We also involve test time augmentation to improve the models interpretability. Finally, no model provides a one-size-fits-all solution for a feasible interpretable model. ConvNext series are suitable for in-distribution samples, while ViT and its variants are suitable for interpreting out-of-distribution samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。