arXiv:2512.19663cs.CVcs.AI2025-12被引 1

融合医学知识的多模态模型显著提升眼底图像与文本的对齐能力。

Beyond CLIP: Knowledge-Enhanced Multimodal Transformers for Cross-Modal Alignment in Diabetic Retinopathy Diagnosis

  • 用视觉、临床文本和结构化数据三路编码,通过联合变压器融合特征。
  • 文本到图像检索召回率99.94%,远超微调版CLIP的1.29%。
  • 适合眼科诊断系统研发者,尤其关注跨模态对齐与零样本泛化。

糖尿病视网膜病变(DR)是全球可预防性失明的主要原因,亟需精准的自动化诊断系统。尽管通用视觉语言模型如对比语言-图像预训练(CLIP)在自然图像任务中表现良好,但在医学领域尤其是眼科图像与文本的跨模态检索中效果有限。本文提出一种新型知识增强型联合嵌入框架,通过多模态变压器架构整合眼底彩照、临床文本及结构化患者数据,填补医学图像-文本对齐的空白。采用ViT-B/16处理眼底图像,Bio-ClinicalBERT解析临床描述,全连接网络处理人口统计与临床特征,三路信息通过带有模态特异性嵌入的联合变压器融合。训练包含多任务目标:模态对间的对比损失、图像与文本重建损失,以及基于ICDR和SDRG分级体系的分类损失。在巴西多标签眼科数据集(BRSET)上的实验表明,本框架实现近乎完美的文本到图像检索性能,召回率@1达99.94%,显著优于微调版CLIP的1.29%;同时保持顶尖分类精度,SDRG为97.05%,ICDR为97.97%。在未见数据集DeepEyeNet的零样本评估中,召回率@1达93.95%,而微调版CLIP仅为0.22%。结果表明,该多模态训练方法有效捕捉医学领域的跨模态关联,兼具卓越检索能力与稳健诊断表现。

原文摘要 · Abstract (English)

Diabetic retinopathy (DR) is a leading cause of preventable blindness worldwide, demanding accurate automated diagnostic systems. While general-domain vision-language models like Contrastive Language-Image Pre-Training (CLIP) perform well on natural image tasks, they struggle in medical domain applications, particularly in cross-modal retrieval for ophthalmological images. We propose a novel knowledge-enhanced joint embedding framework that integrates retinal fundus images, clinical text, and structured patient data through a multimodal transformer architecture to address the critical gap in medical image-text alignment. Our approach employs separate encoders for each modality: a Vision Transformer (ViT-B/16) for retinal images, Bio-ClinicalBERT for clinical narratives, and a multilayer perceptron for structured demographic and clinical features. These modalities are fused through a joint transformer with modality-specific embeddings, trained using multiple objectives including contrastive losses between modality pairs, reconstruction losses for images and text, and classification losses for DR severity grading according to ICDR and SDRG schemes. Experimental results on the Brazilian Multilabel Ophthalmological Dataset (BRSET) demonstrate significant improvements over baseline models. Our framework achieves near-perfect text-to-image retrieval performance with Recall@1 of 99.94% compared to fine-tuned CLIP's 1.29%, while maintaining state-of-the-art classification accuracy of 97.05% for SDRG and 97.97% for ICDR. Furthermore, zero-shot evaluation on the unseen DeepEyeNet dataset validates strong generalizability with 93.95% Recall@1 versus 0.22% for fine-tuned CLIP. These results demonstrate that our multimodal training approach effectively captures cross-modal relationships in the medical domain, establishing both superior retrieval capabilities and robust diagnostic performance.

医学多模态跨模态对齐糖尿病视网膜病变视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。