arXiv:2507.10589eess.IVcs.AI2025-07被引 2

Vision Transformer在儿童肺炎胸片检测中表现优于传统CNN,准确率达88.25%

Comparative Analysis of Vision Transformers and Traditional Deep Learning Approaches for Automated Pneumonia Detection in Chest X-Rays

  • 对比多种模型,ViT架构在肺部影像识别中更优
  • Cross-ViT达88.25%准确率与99.42%召回率
  • 适合医疗影像诊断与快速疫情筛查场景

肺炎,尤其是由新冠等疾病引发的肺炎,仍是全球重大健康挑战,亟需快速精准诊断。本研究系统比较了传统机器学习与先进深度学习方法在胸部X光片(CXRs)上自动检测肺炎的效果。评估涵盖从主成分分析聚类、逻辑回归、支持向量机到改进LeNet、DenseNet-121以及多种Vision Transformer(Deep-ViT、Compact Convolutional Transformer、Cross-ViT)等模型。基于包含5,856张儿童胸片的数据集,结果表明,Vision Transformers尤其在Cross-ViT架构下表现最优,准确率达88.25%,召回率高达99.42%,超越传统CNN。分析显示,模型架构比规模对性能影响更大,7500万参数的Cross-ViT优于更大模型。研究还探讨了计算效率、训练需求及医学诊断中精度与召回率的权衡。结果表明,Vision Transformers为自动化肺炎检测提供了有前景的方向,有望在公共卫生危机中实现更快速准确的诊断。

原文摘要 · Abstract (English)

Pneumonia, particularly when induced by diseases like COVID-19, remains a critical global health challenge requiring rapid and accurate diagnosis. This study presents a comprehensive comparison of traditional machine learning and state-of-the-art deep learning approaches for automated pneumonia detection using chest X-rays (CXRs). We evaluate multiple methodologies, ranging from conventional machine learning techniques (PCA-based clustering, Logistic Regression, and Support Vector Classification) to advanced deep learning architectures including Convolutional Neural Networks (Modified LeNet, DenseNet-121) and various Vision Transformer (ViT) implementations (Deep-ViT, Compact Convolutional Transformer, and Cross-ViT). Using a dataset of 5,856 pediatric CXR images, we demonstrate that Vision Transformers, particularly the Cross-ViT architecture, achieve superior performance with 88.25% accuracy and 99.42% recall, surpassing traditional CNN approaches. Our analysis reveals that architectural choices impact performance more significantly than model size, with Cross-ViT's 75M parameters outperforming larger models. The study also addresses practical considerations including computational efficiency, training requirements, and the critical balance between precision and recall in medical diagnostics. Our findings suggest that Vision Transformers offer a promising direction for automated pneumonia detection, potentially enabling more rapid and accurate diagnosis during health crises.

肺部影像ViT医疗诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。