用双目信息与知识蒸馏,实现高精度可解释的白内障检测
Explainable Deep Learning for Cataract Detection in Retinal Images: A Dual-Eye and Knowledge Distillation Approach
- 融合左右眼图像的双目网络,提升模型判读可靠性
- 轻量级蒸馏模型达98.42%准确率,计算成本大幅降低
- 通过Grad-CAM揭示模型关注医学关键特征,增强临床可信度
白内障是全球致盲的主要原因之一,从眼底图像中早期识别至关重要。本研究基于包含5000名患者左右眼眼底照片的Ocular Disease Recognition数据集,构建深度学习分类流程。对比了CNN、Transformer、轻量级架构及知识蒸馏模型。表现最佳的Swin-Base Transformer模型达到98.58%准确率和0.9836的F1分数。使用Swin-Base知识训练的蒸馏MobileNetV3模型,在显著降低计算开销的前提下,实现98.42%准确率和0.9787 F1分数。提出的蒸馏MobileNet双目变体结合双侧信息,准确率达98.21%。使用Grad-CAM进行可解释性分析显示,模型聚焦于晶状体浑浊、中心模糊等医学相关特征。结果表明,即使采用轻量模型,也可实现高精度且可解释的白内障检测,具备在资源有限地区临床应用潜力。
原文摘要 · Abstract (English)
Cataract remains a leading cause of visual impairment worldwide, and early detection from retinal imaging is critical for timely intervention. We present a deep learning pipeline for cataract classification using the Ocular Disease Recognition dataset, containing left and right fundus photographs from 5000 patients. We evaluated CNNs, transformers, lightweight architectures, and knowledge-distilled models. The top-performing model, Swin-Base Transformer, achieved 98.58% accuracy and an F1-score of 0.9836. A distilled MobileNetV3, trained with Swin-Base knowledge, reached 98.42% accuracy and a 0.9787 F1-score with greatly reduced computational cost. The proposed dual-eye Siamese variant of the distilled MobileNet, integrating information from both eyes, achieved an accuracy of 98.21%. Explainability analysis using Grad-CAM demonstrated that the CNNs concentrated on medically significant features, such as lens opacity and central blur. These results show that accurate, interpretable cataract detection is achievable even with lightweight models, supporting potential clinical integration in resource-limited settings
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。