arXiv:2506.19320cs.CV2025-06中稿 · MICCAI 2025被引 7

让眼底模型持续学习多模态影像,避免遗忘旧知识。

Continual Retinal Vision-Language Pre-training upon Incremental Imaging Modalities

  • 分阶段引入不同眼底影像数据,统一训练视觉语言模型。
  • 在多个模态上表现优于现有方法,遗忘率最低。
  • 适合需要长期更新的眼科AI系统开发者。

传统眼底图像分析模型局限于单一模态任务,忽视了不同成像模态间的互补性,限制了其泛化能力。近年来虽出现视网膜基础模型,但多数仍为模态特定。将多种眼底成像模态整合至统一基础模型具有重要价值。然而,在动态环境中,不同模态的数据常呈增量式到达,需持续预训练。为此,我们提出RetCoP——首个眼底领域持续视觉-语言预训练框架,可增量融合不同成像模态的图像与文本特征。为缓解持续预训练中的灾难性遗忘,我们引入基于代表性图文对的回放策略,以及非对角信息蒸馏方法:前者使模型可重访过往知识,后者显式保持图像与文本表示间的对齐。实验表明,RetCoP在所有对比方法中表现最佳,具备最优泛化能力与最低遗忘率。代码见https://github.com/Yuang-Yao/RetCoP。

原文摘要 · Abstract (English)

Traditional fundus image analysis models focus on single-modal tasks, ignoring fundus modality complementarity, which limits their versatility. Recently, retinal foundation models have emerged, but most still remain modality-specific. Integrating multiple fundus imaging modalities into a single foundation model is valuable. However, in dynamic environments, data from different modalities often arrive incrementally, necessitating continual pre-training. To address this, we propose RetCoP, the first continual vision-language pre-training framework in the fundus domain, which incrementally integrates image and text features from different imaging modalities into a single unified foundation model. To mitigate catastrophic forgetting in continual pre-training, we introduce a rehearsal strategy utilizing representative image-text pairs and an off-diagonal information distillation approach. The former allows the model to revisit knowledge from previous stages, while the latter explicitly preserves the alignment between image and text representations. Experiments show that RetCoP outperforms all the compared methods, achieving the best generalization and lowest forgetting rate. The code can be found at https://github.com/Yuang-Yao/RetCoP.

眼底成像持续学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。