用ViT和可解释AI提升脑病MRI诊断准确率
An Exploratory Approach Towards Investigating and Explaining Vision Transformer and Transfer Learning for Brain Disease Detection
- 对比ViT与多种迁移学习模型在脑病MRI分类中的表现
- ViT达94.39%准确率,优于VGG、ResNet等模型
- 结合GradCAM等方法提升模型可解释性,助医生决策
大脑是管理运动、记忆和思维等重要功能的复杂器官,相关疾病如肿瘤和退行性疾病难以诊断治疗。磁共振成像(MRI)提供高分辨率脑结构图像,是关键诊断工具,但解读复杂。本研究基于孟加拉国数据集,对比Vision Transformer(ViT)与VGG16、VGG19、ResNet50V2、MobileNetV2等迁移学习模型在脑病分类中的表现。ViT擅长捕捉图像全局关系,适合医学影像;迁移学习缓解数据不足问题。同时采用GradCAM、GradCAM++、LayerCAM、ScoreCAM及Faster-ScoreCAM等可解释AI方法解析模型预测。结果表明,ViT分类准确率达94.39%,显著优于其他模型;XAI方法提升了模型透明度,为医生提供精准诊断支持。
原文摘要 · Abstract (English)
The brain is a highly complex organ that manages many important tasks, including movement, memory and thinking. Brain-related conditions, like tumors and degenerative disorders, can be hard to diagnose and treat. Magnetic Resonance Imaging (MRI) serves as a key tool for identifying these conditions, offering high-resolution images of brain structures. Despite this, interpreting MRI scans can be complicated. This study tackles this challenge by conducting a comparative analysis of Vision Transformer (ViT) and Transfer Learning (TL) models such as VGG16, VGG19, Resnet50V2, MobilenetV2 for classifying brain diseases using MRI data from Bangladesh based dataset. ViT, known for their ability to capture global relationships in images, are particularly effective for medical imaging tasks. Transfer learning helps to mitigate data constraints by fine-tuning pre-trained models. Furthermore, Explainable AI (XAI) methods such as GradCAM, GradCAM++, LayerCAM, ScoreCAM, and Faster-ScoreCAM are employed to interpret model predictions. The results demonstrate that ViT surpasses transfer learning models, achieving a classification accuracy of 94.39%. The integration of XAI methods enhances model transparency, offering crucial insights to aid medical professionals in diagnosing brain diseases with greater precision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。