arXiv:2508.14151eess.IVcs.AI2025-08

用深度学习+可解释AI自动定位膝关节MRI病灶,效果优于纯模型方法。

A Systematic Study of Deep Learning Models and xAI Methods for Region-of-Interest Detection in MRI Scans

  • 对比多种模型与可解释技术,发现残差网络在定位上最优。
  • 在MRNet数据集上,ResNet50分类AUC表现最佳,超越视觉变换器。
  • 梯度类激活图(Grad-CAM)提供最清晰的临床可解释性,适合医生信任。

磁共振成像(MRI)是评估膝关节损伤的关键工具,但人工解读耗时且存在观察者差异。本研究系统评估了多种深度学习架构与可解释AI(xAI)技术在膝关节MRI中感兴趣区域(ROI)检测中的表现。采用监督与自监督方法,包括ResNet50、InceptionV3、视觉变换器(ViT)及多类带多层感知机(MLP)的U-Net变体。为提升可解释性,引入Grad-CAM和显著性图等xAI方法。通过分类的AUC以及重建质量的PSNR/SSIM进行评估,并结合定性可视化分析。结果表明,ResNet50在分类与ROI识别中持续领先,优于基于变换器的模型,尤其在有限的MRNet数据集下。混合型U-Net+MLP虽在空间特征利用与解释性方面具潜力,但分类性能仍较低。所有模型中,Grad-CAM提供最符合临床需求的解释。总体而言,基于卷积神经网络的迁移学习在此数据集上最为有效,而更大规模预训练可能释放变换器的潜力。

原文摘要 · Abstract (English)

Magnetic Resonance Imaging (MRI) is an essential diagnostic tool for assessing knee injuries. However, manual interpretation of MRI slices remains time-consuming and prone to inter-observer variability. This study presents a systematic evaluation of various deep learning architectures combined with explainable AI (xAI) techniques for automated region of interest (ROI) detection in knee MRI scans. We investigate both supervised and self-supervised approaches, including ResNet50, InceptionV3, Vision Transformers (ViT), and multiple U-Net variants augmented with multi-layer perceptron (MLP) classifiers. To enhance interpretability and clinical relevance, we integrate xAI methods such as Grad-CAM and Saliency Maps. Model performance is assessed using AUC for classification and PSNR/SSIM for reconstruction quality, along with qualitative ROI visualizations. Our results demonstrate that ResNet50 consistently excels in classification and ROI identification, outperforming transformer-based models under the constraints of the MRNet dataset. While hybrid U-Net + MLP approaches show potential for leveraging spatial features in reconstruction and interpretability, their classification performance remains lower. Grad-CAM consistently provided the most clinically meaningful explanations across architectures. Overall, CNN-based transfer learning emerges as the most effective approach for this dataset, while future work with larger-scale pretraining may better unlock the potential of transformer models.

医学影像深度学习可解释AIMRI分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。