用混合模型提升眼底图像多疾病分类准确率
Retinal Fundus Multi-Disease Image Classification using Hybrid CNN-Transformer-Ensemble Architectures
- 融合CNN、Transformer与集成学习,分层处理眼底图像
- 最高准确率达0.9166,优于基线模型(0.9)
- 适合医疗资源匮乏地区的眼病辅助诊断
本研究针对全球大量人群受视网膜疾病影响但专业医疗资源不足的问题,尤其在非城市地区更为突出。目标是仅通过眼底图像构建一个全面的诊断系统,实现对20种视网膜疾病的精准预测。面对数据集有限、多样性差及类别分布不均的挑战,我们提出创新策略:采用深度卷积神经网络(CNN)、Transformer编码器与集成架构的混合模型,以串行和并行方式结合,提升分类性能。实验表明,C-Tran集成模型表现最优,模型得分达0.9166,显著超越基线模型的0.9。同时,IEViT模型在计算效率上也展现出良好潜力。我们还验证了动态分块提取与领域知识融合在计算机视觉任务中的有效性。研究旨在为偏远地区提供可及性高的眼科诊断解决方案,推动更全面、准确的疾病预测。
原文摘要 · Abstract (English)
Our research is motivated by the urgent global issue of a large population affected by retinal diseases, which are evenly distributed but underserved by specialized medical expertise, particularly in non-urban areas. Our primary objective is to bridge this healthcare gap by developing a comprehensive diagnostic system capable of accurately predicting retinal diseases solely from fundus images. However, we faced significant challenges due to limited, diverse datasets and imbalanced class distributions. To overcome these issues, we have devised innovative strategies. Our research introduces novel approaches, utilizing hybrid models combining deeper Convolutional Neural Networks (CNNs), Transformer encoders, and ensemble architectures sequentially and in parallel to classify retinal fundus images into 20 disease labels. Our overarching goal is to assess these advanced models' potential in practical applications, with a strong focus on enhancing retinal disease diagnosis accuracy across a broader spectrum of conditions. Importantly, our efforts have surpassed baseline model results, with the C-Tran ensemble model emerging as the leader, achieving a remarkable model score of 0.9166, surpassing the baseline score of 0.9. Additionally, experiments with the IEViT model showcased equally promising outcomes with improved computational efficiency. We've also demonstrated the effectiveness of dynamic patch extraction and the integration of domain knowledge in computer vision tasks. In summary, our research strives to contribute significantly to retinal disease diagnosis, addressing the critical need for accessible healthcare solutions in underserved regions while aiming for comprehensive and accurate disease prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。