arXiv:2507.09630cs.CVcs.AI2025-07被引 13

用Transformer模型和可解释AI提升脑卒中CT影像分类准确率

Brain Stroke Detection and Classification Using CT Imaging with Transformer Models and Explainable AI

  • 采用MaxViT等Transformer模型进行多类卒中分类
  • 数据增强后准确率与F1值达98.00%,优于其他模型
  • 结合Grad-CAM++实现决策可视化,适合临床部署

脑卒中是全球主要致死病因之一,早期精准诊断对改善预后至关重要,尤其在急诊环境中及时干预尤为关键。计算机断层扫描(CT)因其快速、易获取且成本低成为首选影像方式。本研究提出一种基于人工智能的多类卒中分类框架(缺血性、出血性、无卒中),使用土耳其卫生部提供的数据集,采用MaxViT等先进视觉Transformer模型进行图像分类,并对比了Vision Transformer、Transformer-in-Transformer和ConvNext等变体。为提升模型泛化能力并缓解类别不平衡问题,引入了包括合成图像生成在内的数据增强技术。经增强训练的MaxViT模型表现最优,准确率和F1-score均达到98.00%,显著优于其他模型及基线方法。研究核心目标是在高精度区分卒中类型的同时,解决AI模型的透明性与可信度问题。为此,引入可解释人工智能(XAI)技术,特别是Grad-CAM++,通过可视化提示模型决策的关键区域,实现对脑卒中病灶的精准定位,构建出可解释、临床可用的早期卒中检测方案。该研究推动了可信AI辅助诊断工具的发展,有助于其在急诊科落地应用,提升卒中患者及时、最优诊疗水平,挽救更多生命。

原文摘要 · Abstract (English)

Stroke is one of the leading causes of death globally, making early and accurate diagnosis essential for improving patient outcomes, particularly in emergency settings where timely intervention is critical. CT scans are the key imaging modality because of their speed, accessibility, and cost-effectiveness. This study proposed an artificial intelligence framework for multiclass stroke classification (ischemic, hemorrhagic, and no stroke) using CT scan images from a dataset provided by the Republic of Turkey's Ministry of Health. The proposed method adopted MaxViT, a state-of-the-art Vision Transformer, as the primary deep learning model for image-based stroke classification, with additional transformer variants (vision transformer, transformer-in-transformer, and ConvNext). To enhance model generalization and address class imbalance, we applied data augmentation techniques, including synthetic image generation. The MaxViT model trained with augmentation achieved the best performance, reaching an accuracy and F1-score of 98.00%, outperforming all other evaluated models and the baseline methods. The primary goal of this study was to distinguish between stroke types with high accuracy while addressing crucial issues of transparency and trust in artificial intelligence models. To achieve this, Explainable Artificial Intelligence (XAI) was integrated into the framework, particularly Grad-CAM++. It provides visual explanations of the model's decisions by highlighting relevant stroke regions in the CT scans and establishing an accurate, interpretable, and clinically applicable solution for early stroke detection. This research contributed to the development of a trustworthy AI-assisted diagnostic tool for stroke, facilitating its integration into clinical practice and enhancing access to timely and optimal stroke diagnosis in emergency departments, thereby saving more lives.

脑卒中CT影像Transformer可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。