用高效模型+注意力机制,实现高准确率且可解释的脑肿瘤MRI自动诊断。
MRI-Based Brain Tumor Detection through an Explainable EfficientNetV2 and MLP-Mixer-Attention Architecture
- 融合EfficientNetV2与注意力MLP-Mixer,提升特征提取能力。
- 在3064张MRI图像上达99.5%准确率,超越已有研究。
- 通过Grad-CAM可视化,明确模型关注病灶区域,适合临床使用。
脑肿瘤是高致死率的严重健康问题,需早期诊断。基于磁共振成像(MRI)的肿瘤判读依赖专家经验,易出错,因此自动化诊断系统需求日益增长。本文提出一种鲁棒且可解释的深度学习模型用于脑肿瘤分类。采用公开的Figshare数据集,包含3,064张T1加权对比增强脑部MRI图像,涵盖三种肿瘤类型。首先评估了九种主流CNN架构,发现EfficientNetV2表现最优,作为主干网络。随后引入基于注意力的MLP-Mixer结构以增强分类能力。通过五折交叉验证,该模型在准确率、精确率、召回率和F1分数上分别达到99.50%、99.47%、99.52%和99.49%,显著优于文献中已有方法。结合Grad-CAM可视化分析,模型能有效聚焦图像中相关病灶区域,提升决策可解释性与临床可靠性。结果表明,EfficientNetV2与注意力MLP-Mixer的结合为临床决策支持系统提供了高精度、可解释的脑肿瘤分类方案。
原文摘要 · Abstract (English)
Brain tumors are serious health problems that require early diagnosis due to their high mortality rates. Diagnosing tumors by examining Magnetic Resonance Imaging (MRI) images is a process that requires expertise and is prone to error. Therefore, the need for automated diagnosis systems is increasing day by day. In this context, a robust and explainable Deep Learning (DL) model for the classification of brain tumors is proposed. In this study, a publicly available Figshare dataset containing 3,064 T1-weighted contrast-enhanced brain MRI images of three tumor types was used. First, the classification performance of nine well-known CNN architectures was evaluated to determine the most effective backbone. Among these, EfficientNetV2 demonstrated the best performance and was selected as the backbone for further development. Subsequently, an attention-based MLP-Mixer architecture was integrated into EfficientNetV2 to enhance its classification capability. The performance of the final model was comprehensively compared with basic CNNs and the methods in the literature. Additionally, Grad-CAM visualization was used to interpret and validate the decision-making process of the proposed model. The proposed model's performance was evaluated using the five-fold cross-validation method. The proposed model demonstrated superior performance with 99.50% accuracy, 99.47% precision, 99.52% recall and 99.49% F1 score. The results obtained show that the model outperforms the studies in the literature. Moreover, Grad-CAM visualizations demonstrate that the model effectively focuses on relevant regions of MRI images, thus improving interpretability and clinical reliability. A robust deep learning model for clinical decision support systems has been obtained by combining EfficientNetV2 and attention-based MLP-Mixer, providing high accuracy and interpretability in brain tumor classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。