融合眼底图像与时间序列数据,实现糖尿病视网膜病变精准早期诊断
Temporal-Enhanced Interpretable Multi-Modal Prognosis and Risk Stratification Framework for Diabetic Retinopathy (TIMM-ProRS)
- 用ViT、CNN和GNN融合多模态数据,捕捉图像与时间动态
- 在多个数据集上达97.8%准确率和0.96 F1-score,优于现有方法
- 结果可解释,适合临床辅助与远程医疗推广
糖尿病视网膜病变(DR)影响全球数百万患者,未来患病率将持续上升,严重威胁视力并加重医疗系统负担。其诊断复杂源于症状与老年黄斑变性、高血压性视网膜病变等重叠,且在资源匮乏地区误诊率高。本文提出TIMM-ProRS框架,结合视觉变压器(ViT)、卷积神经网络(CNN)与图神经网络(GNN),融合眼底图像与时间生物标志物(如HbA1c、视网膜厚度),以捕捉多模态与时间动态特征。在APTOS 2019(训练)、Messidor-2、RFMiD、EyePACS和Messidor-1(验证)等多个数据集上评估,模型达到97.8%准确率与0.96 F1-score,性能优于RSG-Net和DeepDR等现有方法。该方法支持早期、精确、可解释的诊断,助力远程医疗规模化管理,提升全球眼健康可持续性。
原文摘要 · Abstract (English)
Diabetic retinopathy (DR), affecting millions globally with projections indicating a significant rise, poses a severe blindness risk and strains healthcare systems. Diagnostic complexity arises from visual symptom overlap with conditions like age-related macular degeneration and hypertensive retinopathy, exacerbated by high misdiagnosis rates in underserved regions. This study introduces TIMM-ProRS, a novel deep learning framework integrating Vision Transformer (ViT), Convolutional Neural Network (CNN), and Graph Neural Network (GNN) with multi-modal fusion. TIMM-ProRS uniquely leverages both retinal images and temporal biomarkers (HbA1c, retinal thickness) to capture multi-modal and temporal dynamics. Evaluated comprehensively across diverse datasets including APTOS 2019 (trained), Messidor-2, RFMiD, EyePACS, and Messidor-1 (validated), the model achieves 97.8\% accuracy and an F1-score of 0.96, demonstrating state-of-the-art performance and outperforming existing methods like RSG-Net and DeepDR. This approach enables early, precise, and interpretable diagnosis, supporting scalable telemedical management and enhancing global eye health sustainability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。