arXiv:2507.04999cs.CV2025-07被引 1

用最优传输对齐眼科影像,缺模态也能准诊断。

Robust Incomplete-Modality Alignment for Ophthalmic Disease Grading and Diagnosis via Labeled Optimal Transport

  • 基于最优传输实现多尺度跨模态特征对齐,精准匹配病变语义。
  • 在三种数据集上缺模态场景下均达当前最佳,完整数据也领先。
  • 适合医疗影像分析、弱监督学习方向的研究者参考。

基于多模态眼底影像的诊断融合彩色眼底图与光学相干断层扫描(OCT),以全面评估眼部病变。然而,医疗资源分布不均导致临床中常出现模态缺失,严重降低诊断准确性。现有方法如模态补全和知识蒸馏存在局限:补全难以重建局灶性病变特征,且眼底图风格差异大;蒸馏依赖完全配对数据。为此,我们提出一种鲁棒的多模态对齐与融合框架,可有效应对模态缺失。针对OCT与眼底图特征差异,强调同类别语义对齐,并显式学习模态间软匹配,使缺失模态能利用已有信息实现稳健跨模态对齐。具体采用多尺度最优传输:通过预测类别原型实现类别级对齐,通过跨模态共享特征传输实现特征级对齐。同时设计非对称融合策略,充分挖掘OCT与眼底图的差异化优势。在三个大规模眼科多模态数据集上的实验表明,模型在多种模态缺失场景下均表现优异,完整模态与跨模态缺失条件下均达到当前最优性能。代码已开源。

原文摘要 · Abstract (English)

Multimodal ophthalmic imaging-based diagnosis integrates color fundus image with optical coherence tomography (OCT) to provide a comprehensive view of ocular pathologies. However, the uneven global distribution of healthcare resources often results in real-world clinical scenarios encountering incomplete multimodal data, which significantly compromises diagnostic accuracy. Existing commonly used pipelines, such as modality imputation and distillation methods, face notable limitations: 1)Imputation methods struggle with accurately reconstructing key lesion features, since OCT lesions are localized, while fundus images vary in style. 2)distillation methods rely heavily on fully paired multimodal training data. To address these challenges, we propose a novel multimodal alignment and fusion framework capable of robustly handling missing modalities in the task of ophthalmic diagnostics. By considering the distinctive feature characteristics of OCT and fundus images, we emphasize the alignment of semantic features within the same category and explicitly learn soft matching between modalities, allowing the missing modality to utilize existing modality information, achieving robust cross-modal feature alignment under the missing modality. Specifically, we leverage the Optimal Transport for multi-scale modality feature alignment: class-wise alignment through predicted class prototypes and feature-wise alignment via cross-modal shared feature transport. Furthermore, we propose an asymmetric fusion strategy that effectively exploits the distinct characteristics of OCT and fundus modalities. Extensive evaluations on three large ophthalmic multimodal datasets demonstrate our model's superior performance under various modality-incomplete scenarios, achieving Sota performance in both complete modality and inter-modality incompleteness conditions. Code is available at https://github.com/Qinkaiyu/RIMA

医学影像多模态最优传输缺模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。