双目眼底图像分类新模型,提升对称性病灶检测准确率
DMS-Net:Dual-Modal Multi-Scale Siamese Network for Binocular Fundus Image Classification
- 采用双模态多尺度孪生网络,同步处理左右眼图像特征
- 在ODIR-5K数据集上达82.9%准确率、84.5%召回率
- 适合眼科疾病辅助诊断与医疗机器人场景应用
眼科疾病带来重大全球健康负担。传统诊断方法及单眼图像深度学习模型常忽略双眼病理关联。实际医疗机器人诊断中,成对眼底图像(双目眼底图像)是重要诊断依据。为此,我们提出DMS-Net——一种用于双目眼底图像分类的双模态多尺度孪生网络。该框架采用共享权重的孪生ResNet-152结构,同步提取双侧眼底图像深层语义特征。为应对病灶边界模糊和病理分布扩散等挑战,引入全向空间整合模块(OSIM),通过多尺度自适应池化与空间注意力机制实现多分辨率特征聚合。同时,设计校准式相似语义融合模块(CASFM),利用空间-语义重校准与双向注意力增强跨模态交互,聚合非模态眼底结构表示。为进一步挖掘双侧图像中病灶差异语义信息,提出跨模态对比对齐模块(CCAM);为强化病灶相关语义聚合,引入跨模态集成对齐模块(CIAM)。在ODIR-5K数据集上的评估表明,DMS-Net达到82.9%准确率、84.5%召回率与83.2%的Cohen's kappa系数,展现出检测对称性病灶的强能力,并有助于提升眼科疾病临床决策水平。代码与处理后数据集将后续发布。
原文摘要 · Abstract (English)
Ophthalmic diseases pose a significant global health burden. However, traditional diagnostic methods and existing monocular image-based deep learning approaches often overlook the pathological correlations between the two eyes. In practical medical robotic diagnostic scenarios, paired retinal images (binocular fundus images) are frequently required as diagnostic evidence. To address this, we propose DMS-Net-a dual-modal multi-scale siamese network for binocular retinal image classification. The framework employs a weight-sharing siamese ResNet-152 architecture to concurrently extract deep semantic features from bilateral fundus images. To tackle challenges like indistinct lesion boundaries and diffuse pathological distributions, we introduce the OmniPool Spatial Integrator Module (OSIM), which achieves multi-resolution feature aggregation through multi-scale adaptive pooling and spatial attention mechanisms. Furthermore, the Calibrated Analogous Semantic Fusion Module (CASFM) leverages spatial-semantic recalibration and bidirectional attention mechanisms to enhance cross-modal interaction, aggregating modality-agnostic representations of fundus structures. To fully exploit the differential semantic information of lesions present in bilateral fundus features, we introduce the Cross-Modal Contrastive Alignment Module (CCAM). Additionally, to enhance the aggregation of lesion-correlated semantic information, we introduce the Cross-Modal Integrative Alignment Module (CIAM). Evaluation on the ODIR-5K dataset demonstrates that DMS-Net achieves state-of-the-art performance with an accuracy of 82.9%, recall of 84.5%, and a Cohen's kappa coefficient of 83.2%, showcasing robust capacity in detecting symmetrical pathologies and improving clinical decision-making for ocular diseases. Code and the processed dataset will be released subsequently.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。