用知识蒸馏提升胃肠道疾病诊断准确率与可解释性
A Graph-Augmented knowledge Distillation based Dual-Stream Vision Transformer with Region-Aware Attention for Gastrointestinal Disease Classification with Explainable AI
- 双流架构融合Swin与ViT优势,通过知识蒸馏让小模型继承大模型能力
- 在两个数据集上准确率达99.78%和99.28%,平均AUC达1.0000
- 模型预测聚焦病灶区域,适合临床部署的高效可解释诊断系统
从内窥镜与病理图像中准确分类胃肠道疾病仍是医学诊断的重大挑战,主要源于数据量庞大及类间视觉差异细微。本研究提出一种基于知识蒸馏的双流深度学习框架,其中高容量教师模型结合Swin Transformer的全局上下文推理能力与Vision Transformer的局部细粒度特征提取能力。学生网络采用紧凑的Tiny-ViT结构,通过软标签蒸馏继承教师的语义与形态知识,在效率与诊断准确性间取得平衡。使用两个精心构建的无线胶囊内镜数据集,涵盖主要胃肠道疾病类别,确保样本均衡并减少偏差。所提框架在数据集1和数据集2上分别达到0.9978和0.9928的准确率,平均AUC为1.0000,表明近乎完美的判别能力。通过Grad-CAM、LIME和Score-CAM的可解释性分析证实,模型预测基于临床相关的组织区域与病理形态线索,验证了其透明性与可靠性。Tiny-ViT在计算复杂度显著降低的前提下,实现与基于Transformer的教师模型相当的诊断性能,并具备更快推理速度,适用于资源受限的临床环境。总体而言,该框架为人工智能辅助胃肠道疾病诊断提供了一种鲁棒、可解释且可扩展的解决方案,推动未来兼容临床实践的智能内镜筛查发展。
原文摘要 · Abstract (English)
The accurate classification of gastrointestinal diseases from endoscopic and histopathological imagery remains a significant challenge in medical diagnostics, mainly due to the vast data volume and subtle variation in inter-class visuals. This study presents a hybrid dual-stream deep learning framework built on teacher-student knowledge distillation, where a high-capacity teacher model integrates the global contextual reasoning of a Swin Transformer with the local fine-grained feature extraction of a Vision Transformer. The student network was implemented as a compact Tiny-ViT structure that inherits the teacher's semantic and morphological knowledge via soft-label distillation, achieving a balance between efficiency and diagnostic accuracy. Two carefully curated Wireless Capsule Endoscopy datasets, encompassing major GI disease classes, were employed to ensure balanced representation and prevent inter-sample bias. The proposed framework achieved remarkable performance with accuracies of 0.9978 and 0.9928 on Dataset 1 and Dataset 2 respectively, and an average AUC of 1.0000, signifying near-perfect discriminative capability. Interpretability analyses using Grad-CAM, LIME, and Score-CAM confirmed that the model's predictions were grounded in clinically significant tissue regions and pathologically relevant morphological cues, validating the framework's transparency and reliability. The Tiny-ViT demonstrated diagnostic performance with reduced computational complexity comparable to its transformer-based teacher while delivering faster inference, making it suitable for resource-constrained clinical environments. Overall, the proposed framework provides a robust, interpretable, and scalable solution for AI-assisted GI disease diagnosis, paving the way toward future intelligent endoscopic screening that is compatible with clinical practicality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。