用注意力机制融合OCT与OCTA图像,提升糖尿病视网膜病变诊断准确率。
Cross-Modal Fusion of OCT and OCT angiography enface for Improved Diagnostics of Diabetic Retinopathy

- 采用双向跨模态注意力网络融合OCT结构图与OCTA血流图
- 使用生成的OCTA图像在多个数据集上达到接近真实OCTA的性能
- 计算生成的OCTA可降低设备成本,适合资源有限地区推广
糖尿病视网膜病变(DR)是全球主要致盲原因,亟需精准且易获取的筛查工具。光学相干断层扫描(OCT)提供高分辨率视网膜结构信息,而OCT血管成像(OCTA)则补充了对DR诊断至关重要的血管信息。本研究提出一种基于双向跨模态注意力网络的OCT B-scan与单通道OCTA俯视图融合方法,用于自动化DR分类。在包含730名受试者的OCT500和UIC两个独立数据集上,评估了模型在同数据集、联合数据集及跨数据集泛化下的表现。以仅使用OCT图像训练的ConvNeXt V2模型为单模态基线。除真实OCTA(GT OCTA)外,还探索了从OCT图像生成的转换OCTA(TR OCTA),无需专用OCTA硬件。实验表明,跨模态融合在所有评估场景中均优于单模态OCT分类。融合真实OCTA显著提升分类准确率与判别能力,而转换OCTA在多数场景中达到相当或更优效果。此外,转换OCTA提升了敏感性与跨数据集泛化能力,表明对域偏移具有更强鲁棒性。结果表明,基于注意力的OCT-OCTA俯视图融合能为DR检测带来临床意义的改进,并证实计算生成的OCTA可作为低成本实用替代方案,推动高性能视网膜筛查系统在资源受限环境中的广泛应用。
原文摘要 · Abstract (English)
Diabetic retinopathy (DR) is a leading cause of vision impairment worldwide, highlighting the need for accurate and accessible screening tools. Optical Coherence Tomography (OCT) provides high-resolution structural information of the retina, whereas OCT angiography (OCTA) offers complementary vascular information that is highly relevant for DR diagnosis. In this study, we propose a cross-modal fusion of OCT B-scans with single-channel en face OCTA using a bidirectional cross-modal attention network for automated DR classification. Two independent datasets, OCT500 and UIC, comprising 730 subjects in total, were utilized to evaluate performance under within-dataset, combined-dataset, and cross-dataset generalization settings. A ConvNeXt V2 model trained solely on OCT images served as the unimodal baseline. In addition to ground-truth (GT) OCTA, we explored the use of translated (TR) OCTA generated from OCT scans, eliminating the requirement for dedicated OCTA hardware. Experimental results demonstrate that cross-modal fusion consistently outperforms unimodal OCT classification across all evaluation scenarios. Fusion with GT OCTA improved classification accuracy and discriminative performance, while TR OCTA achieved comparable or superior results in most settings. Furthermore, TR OCTA improved sensitivity and cross-dataset generalization, indicating enhanced robustness to domain shifts. These findings demonstrate that attention-based OCT-OCTA en face fusion provides clinically meaningful improvements for DR detection and suggest that computationally generated OCTA can serve as a practical, low-cost alternative to hardware-acquired OCTA, enabling broader deployment of high-performance retinal screening systems in resource-limited clinical environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。