融合彩色与红外眼底图像,提升糖尿病视网膜病变分级准确率
Cross-Fundus Transformer for Multi-modal Diabetic Retinopathy Grading with Cataract
- 构建双流Transformer架构,通过跨模态注意力融合两种眼底图像特征
- 在1713对眼底图像上实现更优分级性能,验证多模态优势
- 适合眼科医学AI研究者,尤其关注多模态影像融合的场景
糖尿病视网膜病变(DR)是全球致盲的主要原因,也是糖尿病的常见并发症。彩色眼底照相(CFP)和红外眼底照相(IFP)作为两种不同的DR评估影像工具,在临床上高度相关且互补。据我们所知,这是首个探索将CFP与IFP信息融合以提升DR分级准确性的多模态深度学习框架。本文提出双流结构Cross-Fundus Transformer(CFT),利用ViT提取两种模态的特征,并引入精心设计的跨眼底注意力(CFA)模块,捕捉CFP与IFP之间的对应关系。此外,采用单模态与多模态联合监督策略以最大化分级性能。在包含1713对多模态眼底图像的临床数据集上进行了大量实验,结果证明了该方法的优越性。代码将公开共享。
原文摘要 · Abstract (English)
Diabetic retinopathy (DR) is a leading cause of blindness worldwide and a common complication of diabetes. As two different imaging tools for DR grading, color fundus photography (CFP) and infrared fundus photography (IFP) are highly-correlated and complementary in clinical applications. To the best of our knowledge, this is the first study that explores a novel multi-modal deep learning framework to fuse the information from CFP and IFP towards more accurate DR grading. Specifically, we construct a dual-stream architecture Cross-Fundus Transformer (CFT) to fuse the ViT-based features of two fundus image modalities. In particular, a meticulously engineered Cross-Fundus Attention (CFA) module is introduced to capture the correspondence between CFP and IFP images. Moreover, we adopt both the single-modality and multi-modality supervisions to maximize the overall performance for DR grading. Extensive experiments on a clinical dataset consisting of 1,713 pairs of multi-modal fundus images demonstrate the superiority of our proposed method. Our code will be released for public access.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。