arXiv:2507.21156eess.IVcs.CV2025-07被引 10

对比卷积网络与视觉变换器在医学图像分类中的表现,指导临床AI模型选型。

Comparative Analysis of Vision Transformers and Convolutional Neural Networks for Medical Image Classification

  • 用四种主流模型在三种医学任务上做横向对比,覆盖胸部X光、脑瘤和皮肤癌分类。
  • ResNet-50在肺炎检测中达98.37%准确率,DeiT-Small在脑瘤分类中表现最佳(92.16%)。
  • 结果表明模型性能高度依赖任务类型,为临床AI系统设计提供关键参考。

视觉变换器(ViTs)的兴起彻底改变了计算机视觉领域,但在医学影像中的有效性与传统卷积神经网络(CNNs)相比仍缺乏深入研究。本研究对CNN与ViT架构在三个关键医学影像任务中进行了全面比较:胸部X光肺炎检测、脑肿瘤分类和皮肤癌黑色素瘤检测。评估了四种先进模型——ResNet-50、EfficientNet-B0、ViT-Base和DeiT-Small——在总计8,469张医学图像的数据集上的表现。结果显示不同任务存在模型优势差异:ResNet-50在胸部X光分类中达到98.37%准确率,DeiT-Small在脑肿瘤检测中表现最优,准确率达92.16%,而EfficientNet-B0在皮肤癌分类中以81.84%的准确率领先。这些发现为医疗AI应用中的模型选择提供了重要依据,强调了在临床决策支持系统中进行任务特定架构选型的重要性。

原文摘要 · Abstract (English)

The emergence of Vision Transformers (ViTs) has revolutionized computer vision, yet their effectiveness compared to traditional Convolutional Neural Networks (CNNs) in medical imaging remains under-explored. This study presents a comprehensive comparative analysis of CNN and ViT architectures across three critical medical imaging tasks: chest X-ray pneumonia detection, brain tumor classification, and skin cancer melanoma detection. We evaluated four state-of-the-art models - ResNet-50, EfficientNet-B0, ViT-Base, and DeiT-Small - across datasets totaling 8,469 medical images. Our results demonstrate task-specific model advantages: ResNet-50 achieved 98.37% accuracy on chest X-ray classification, DeiT-Small excelled at brain tumor detection with 92.16% accuracy, and EfficientNet-B0 led skin cancer classification at 81.84% accuracy. These findings provide crucial insights for practitioners selecting architectures for medical AI applications, highlighting the importance of task-specific architecture selection in clinical decision support systems.

医学影像视觉变换器模型对比分类任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。