arXiv:2509.08624cs.CVcs.AI2025-09被引 2

用无配对OCT数据学习病变部位先验,提升眼底图诊断准确率。

UOPSL: Unpaired OCT Predilection Sites Learning for Fundus Image Diagnosis Augmentation

  • 通过对比学习从无配对图像中提取OCT空间先验
  • 在9个数据集28类疾病上显著优于现有方法
  • 适合缺乏配对OCT数据的临床场景使用

近年来,人工智能驱动的多模态医学影像诊断在眼科疾病识别方面取得显著进展。然而,获取配对的眼底与OCT图像成本高昂。虽然眼底摄影简单且经济,但OCT数据稀缺及模态不平衡限制了进一步发展。传统仅依赖眼底或文本特征的方法难以捕捉病灶定位的细粒度空间信息,因各成像模态提供不同的病变偏好线索。本研究提出一种新型无配对多模态框架UOPSL,利用大量OCT-derived空间先验动态识别病变偏好部位,增强基于眼底图像的疾病识别能力。通过扩展疾病文本描述,桥梁式连接无配对的眼底与OCT图像。首先,在大规模无配对眼底和OCT图像上采用对比学习,同时在OCT隐空间中学习病变部位矩阵;经充分优化后,该矩阵捕获了OCT特征空间中的病灶定位模式。在下游分类任务(仅基于眼底图像)的微调或推理阶段,虽无配对OCT数据,仍可去除OCT输入,直接使用病变部位矩阵辅助眼底图像分类学习。在9个不同数据集、28个关键类别上的大量实验表明,本框架超越现有基准。

原文摘要 · Abstract (English)

Significant advancements in AI-driven multimodal medical image diagnosis have led to substantial improvements in ophthalmic disease identification in recent years. However, acquiring paired multimodal ophthalmic images remains prohibitively expensive. While fundus photography is simple and cost-effective, the limited availability of OCT data and inherent modality imbalance hinder further progress. Conventional approaches that rely solely on fundus or textual features often fail to capture fine-grained spatial information, as each imaging modality provides distinct cues about lesion predilection sites. In this study, we propose a novel unpaired multimodal framework \UOPSL that utilizes extensive OCT-derived spatial priors to dynamically identify predilection sites, enhancing fundus image-based disease recognition. Our approach bridges unpaired fundus and OCTs via extended disease text descriptions. Initially, we employ contrastive learning on a large corpus of unpaired OCT and fundus images while simultaneously learning the predilection sites matrix in the OCT latent space. Through extensive optimization, this matrix captures lesion localization patterns within the OCT feature space. During the fine-tuning or inference phase of the downstream classification task based solely on fundus images, where paired OCT data is unavailable, we eliminate OCT input and utilize the predilection sites matrix to assist in fundus image classification learning. Extensive experiments conducted on 9 diverse datasets across 28 critical categories demonstrate that our framework outperforms existing benchmarks.

多模态眼底图像OCT无配对学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。