arXiv:2411.11376eess.IVcs.CV2024-11被引 8

Vision Transformer在肺部X光诊断中表现优于传统CNN,最高准确率达97.83%。

Lung Disease Detection with Vision Transformers: A Comparative Study of Machine Learning Methods

  • 用ViT直接分析完整X光片,无需手动分割肺部区域
  • 全图ViT模型准确率高达97.83%,在八类疾病上AUC达94.54%
  • 结果表明ViT可简化预处理流程,适合临床辅助诊断场景

近年来医疗影像分析主要依赖卷积神经网络(CNN),在胸部X光分类任务中已取得显著成果,如AutoThorax-Net报告的92% AUC和ChexNet实现的88% AUC。然而,在医学领域,微小的准确率提升也可能带来重要临床影响。本研究探讨了视觉变压器(ViT)这一前沿机器学习架构在胸部X光分析中的应用,旨在突破诊断准确率的极限。通过对比两种基于ViT的方法——一种使用完整胸部X光图像,另一种聚焦于分割后的肺部区域,实验表明两者均超越传统CNN模型性能:全图ViT达到97.83%准确率,肺部分割ViT达96.58%准确率;当标签数增至八类时,其AUC为94.54%。全图方法在精度、召回率、F1分数和AUC-ROC等各项指标上均表现更优。结果表明,ViT能有效捕捉胸部X光中的相关特征,无需显式肺部分割,有望简化预处理流程并保持高精度。该研究为变压器架构在医疗影像分析中的有效性提供了有力证据,凸显其在临床诊断中提升精确度的潜力。

原文摘要 · Abstract (English)

Recent advancements in medical image analysis have predominantly relied on Convolutional Neural Networks (CNNs), achieving impressive performance in chest X-ray classification tasks, such as the 92% AUC reported by AutoThorax-Net and the 88% AUC achieved by ChexNet in classifcation tasks. However, in the medical field, even small improvements in accuracy can have significant clinical implications. This study explores the application of Vision Transformers (ViT), a state-of-the-art architecture in machine learning, to chest X-ray analysis, aiming to push the boundaries of diagnostic accuracy. I present a comparative analysis of two ViT-based approaches: one utilizing full chest X-ray images and another focusing on segmented lung regions. Experiments demonstrate that both methods surpass the performance of traditional CNN-based models, with the full-image ViT achieving up to 97.83% accuracy and the lung-segmented ViT reaching 96.58% accuracy in classifcation of diseases on three label and AUC of 94.54% when label numbers are increased to eight. Notably, the full-image approach showed superior performance across all metrics, including precision, recall, F1 score, and AUC-ROC. These findings suggest that Vision Transformers can effectively capture relevant features from chest X-rays without the need for explicit lung segmentation, potentially simplifying the preprocessing pipeline while maintaining high accuracy. This research contributes to the growing body of evidence supporting the efficacy of transformer-based architectures in medical image analysis and highlights their potential to enhance diagnostic precision in clinical settings.

肺部疾病视觉变压器医学影像分类任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。