用优质图文数据训练眼科视觉语言模型,提升跨任务泛化能力
MM-Retinal V2: Transfer an Elite Knowledge Spark into Fundus Vision-Language Pretraining
- 从高质量数据中提取知识,注入公共数据集进行预训练
- 零样本、少样本下表现媲美依赖私有数据的顶尖模型
- 适合医学图像分析与多模态预训练研究者使用
视觉-语言预训练(VLP)在视网膜图像分析中被用于提升下游任务的泛化能力。尽管近期方法已取得良好成果,但其高度依赖大规模私有图文数据,且对预训练方式关注不足,限制了进一步发展。本文提出MM-Retinal V2,一个包含CFP、FFA和OCT模态的高质量图文配对数据集。同时,我们设计了KeepFIT V2模型,通过将精英数据中的知识注入公共数据进行预训练。具体地,先对文本编码器进行初步文本预训练,以引入眼科专业术语知识;再设计混合图文知识注入模块,结合对比学习的全局语义概念与生成学习的局部外观细节实现知识迁移。在零样本、少样本及线性探测设置下的广泛实验表明,KeepFIT V2具有优异的泛化与迁移能力,性能可媲美基于大规模私有图文数据训练的最先进模型。数据集与模型已公开于https://github.com/lxirich/MM-Retinal。
原文摘要 · Abstract (English)
Vision-language pretraining (VLP) has been investigated to generalize across diverse downstream tasks for fundus image analysis. Although recent methods showcase promising achievements, they significantly rely on large-scale private image-text data but pay less attention to the pretraining manner, which limits their further advancements. In this work, we introduce MM-Retinal V2, a high-quality image-text paired dataset comprising CFP, FFA, and OCT image modalities. Then, we propose a novel fundus vision-language pretraining model, namely KeepFIT V2, which is pretrained by integrating knowledge from the elite data spark into categorical public datasets. Specifically, a preliminary textual pretraining is adopted to equip the text encoder with primarily ophthalmic textual knowledge. Moreover, a hybrid image-text knowledge injection module is designed for knowledge transfer, which is essentially based on a combination of global semantic concepts from contrastive learning and local appearance details from generative learning. Extensive experiments across zero-shot, few-shot, and linear probing settings highlight the generalization and transferability of KeepFIT V2, delivering performance competitive to state-of-the-art fundus VLP models trained on large-scale private image-text datasets. Our dataset and model are publicly available via https://github.com/lxirich/MM-Retinal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。