利用多图与临床信息提升眼底病筛查的视觉语言模型
Context-Aware Vision Language Foundation Models for Ocular Disease Screening in Retinal Images
- 融合同一就诊的多张眼底图和患者历史诊疗数据进行联合分析
- 在自建数据集上实现0.851-0.9999的糖尿病视网膜病变分级AUC
- 适合医疗影像领域需应对数据分布变化的研究者使用
基础模型是通过海量多样数据训练得到的通用系统,具备良好的可迁移性,特别适用于医学影像领域。本文基于包含近70万张眼底照片的OPHDIAT数据集,推进视觉语言基础模型(VLF)在眼科疾病筛查中的应用。该数据集涵盖患者糖尿病健康状况、治疗记录、医生文本诊断及每次检查的多张眼底图像。在FLAIR模型基础上,提出新型上下文感知的VLF模型,通过联合分析同一就诊的多张图像或利用过往诊断与上下文信息,充分挖掘数据集丰富性并增强对域偏移的鲁棒性。模型在数据集内部测试子集和公开外部数据集上评估,内部表现优异(DR分级AUC:0.851–0.9999),对外部数据也展现出良好泛化能力(AUC:0.631–0.913)。
原文摘要 · Abstract (English)
Foundation models are large-scale versatile systems trained on vast quantities of diverse data to learn generalizable representations. Their adaptability with minimal fine-tuning makes them particularly promising for medical imaging, where data variability and domain shifts are major challenges. Currently, two types of foundation models dominate the literature: self-supervised models and more recent vision-language models. In this study, we advance the application of vision-language foundation (VLF) models for ocular disease screening using the OPHDIAT dataset, which includes nearly 700,000 fundus photographs from a French diabetic retinopathy (DR) screening network. This dataset provides extensive clinical data (patient-specific information such as diabetic health conditions, and treatments), labeled diagnostics, ophthalmologists text-based findings, and multiple retinal images for each examination. Building on the FLAIR model $\unicode{x2013}$ a VLF model for retinal pathology classification $\unicode{x2013}$ we propose novel context-aware VLF models (e.g jointly analyzing multiple images from the same visit or taking advantage of past diagnoses and contextual data) to fully leverage the richness of the OPHDIAT dataset and enhance robustness to domain shifts. Our approaches were evaluated on both in-domain (a testing subset of OPHDIAT) and out-of-domain data (public datasets) to assess their generalization performance. Our model demonstrated improved in-domain performance for DR grading, achieving an area under the curve (AUC) ranging from 0.851 to 0.9999, and generalized well to ocular disease detection on out-of-domain data (AUC: 0.631-0.913).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。