arXiv:2602.03622cs.CVphysics.med-ph2026-02被引 1

融合眼底造影等多模态数据,提升视网膜疾病诊断准确率。

Quasi-multimodal-based pathophysiological feature learning for retinal disease diagnosis

  • 合成眼底造影、多光谱成像等多源数据,构建统一分析框架。
  • 多标签分类F1-score达0.683,糖尿病视网膜病变分级准确率84.2%。
  • 适合医学影像分析、AI辅助诊断研究者参考。

视网膜疾病可通过多模态数据的互补信号有效识别与诊断。然而,眼科实践中多模态诊断常面临数据异构、潜在侵入性、配准复杂等问题。为此,本文提出一种整合多模态数据生成与融合的统一框架,用于视网膜疾病分类与分级。具体而言,合成数据包含眼底荧光血管造影(FFA)、多光谱成像(MSI)及强调隐匿病灶与视盘/杯区域的显著图。并行模型独立学习各模态特异性表示,捕捉跨病理生理特征。这些特征在模态内与跨模态间自适应校准,根据下游任务实现信息剪枝与灵活融合。通过图像与特征空间的可视化,系统得到充分解释。在两个公开数据集上的实验表明,该方法在多标签分类(F1-score: 0.683,AUC: 0.953)和糖尿病视网膜病变分级(准确率: 0.842,Kappa: 0.861)任务上优于现有先进方法。本工作不仅提升了视网膜疾病筛查的准确性与效率,也为多种医学影像模态的数据增强提供了可扩展框架。

原文摘要 · Abstract (English)

Retinal diseases spanning a broad spectrum can be effectively identified and diagnosed using complementary signals from multimodal data. However, multimodal diagnosis in ophthalmic practice is typically challenged in terms of data heterogeneity, potential invasiveness, registration complexity, and so on. As such, a unified framework that integrates multimodal data synthesis and fusion is proposed for retinal disease classification and grading. Specifically, the synthesized multimodal data incorporates fundus fluorescein angiography (FFA), multispectral imaging (MSI), and saliency maps that emphasize latent lesions as well as optic disc/cup regions. Parallel models are independently trained to learn modality-specific representations that capture cross-pathophysiological signatures. These features are then adaptively calibrated within and across modalities to perform information pruning and flexible integration according to downstream tasks. The proposed learning system is thoroughly interpreted through visualizations in both image and feature spaces. Extensive experiments on two public datasets demonstrated the superiority of our approach over state-of-the-art ones in the tasks of multi-label classification (F1-score: 0.683, AUC: 0.953) and diabetic retinopathy grading (Accuracy:0.842, Kappa: 0.861). This work not only enhances the accuracy and efficiency of retinal disease screening but also offers a scalable framework for data augmentation across various medical imaging modalities.

视网膜疾病多模态学习医学影像数据融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。