EyeCLIP用277万张多模态眼病图像,实现跨模态精准诊断。
EyeCLIP: A visual-language foundation model for multi-modal ophthalmic image analysis
- 融合图像与文本的多模态对比学习,构建共享表征
- 在14个数据集上达到顶尖性能,少样本场景表现优异
- 适合眼科疾病研究者、临床医生及医疗AI开发者使用
早期发现青光眼、黄斑变性、糖尿病视网膜病变等眼病对预防失明至关重要。尽管人工智能基础模型具有巨大潜力,但现有眼科基础模型多局限于单一模态,而眼病诊断需整合多种模态信息。此外,由于眼科疾病呈现长尾分布,传统全监督或无监督学习常难奏效。因此,引入临床文本以覆盖更广疾病谱系尤为关键。我们提出EyeCLIP,基于超过277万张带有部分文本信息的多模态眼科图像构建视觉-语言基础模型。为充分利用大规模有标签与无标签数据,设计结合自监督重建、多模态图像对比学习和图像-文本对比学习的预训练策略,实现多模态共享表征。在14个基准数据集上的评估表明,EyeCLIP可迁移至多种眼病及系统性疾病下游任务,于疾病分类、视觉问答和跨模态检索中均达当前最优表现。尤其在真实世界长尾场景下展现少样本甚至零样本能力,显著优于先前方法。
原文摘要 · Abstract (English)
Early detection of eye diseases like glaucoma, macular degeneration, and diabetic retinopathy is crucial for preventing vision loss. While artificial intelligence (AI) foundation models hold significant promise for addressing these challenges, existing ophthalmic foundation models primarily focus on a single modality, whereas diagnosing eye diseases requires multiple modalities. A critical yet often overlooked aspect is harnessing the multi-view information across various modalities for the same patient. Additionally, due to the long-tail nature of ophthalmic diseases, standard fully supervised or unsupervised learning approaches often struggle. Therefore, it is essential to integrate clinical text to capture a broader spectrum of diseases. We propose EyeCLIP, a visual-language foundation model developed using over 2.77 million multi-modal ophthalmology images with partial text data. To fully leverage the large multi-modal unlabeled and labeled data, we introduced a pretraining strategy that combines self-supervised reconstructions, multi-modal image contrastive learning, and image-text contrastive learning to learn a shared representation of multiple modalities. Through evaluation using 14 benchmark datasets, EyeCLIP can be transferred to a wide range of downstream tasks involving ocular and systemic diseases, achieving state-of-the-art performance in disease classification, visual question answering, and cross-modal retrieval. EyeCLIP represents a significant advancement over previous methods, especially showcasing few-shot, even zero-shot capabilities in real-world long-tail scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。