让眼底影像的量化数据自动转化为有依据的诊断结论。
RetiBridge: Bridging Quantitative Retinal Biomarkers and Qualitative Diagnosis with a Knowledge-Guided Multimodal Large Language Model
- 用多模态大模型融合眼底照片、OCT和文本,建立从数据到诊断的路径。
- 在15611个样本上测试,对31种OCT和6种CFP生物标志物识别准确率领先。
- 适合眼科研究者和临床辅助诊断系统开发者使用。
彩色眼底照相和光学相干断层扫描捕捉的视网膜生物标志物对眼病及全身性疾病具有重要临床价值。多模态大语言模型在眼底图像解读中展现出潜力,但现有眼科模型很少量化这些关键生物标志物,也未将测量结果明确转化为有证据支持的定性诊断结论。为此,我们提出RetiBridge,一种基于知识引导的多模态大语言模型,可联合分析彩色眼底照相(CFP)、光学相干断层扫描(OCT)和文本,明确连接定量视网膜生物标志物与定性临床推断及连贯的诊断结论。RetiBridge结合知识引导指令生成、OCT生物标志物对齐与监督式多模态指令微调,学习从量化指标到定性诊断的路径。基于英国生物银行的15,611对配对的CFP-OCT样本,包含31种OCT和6种CFP生物标志物,我们构建了“有依据的眼科理解”基准,用于评估诊断分类、报告质量及细粒度临床质量。尽管仅使用70亿参数的Qwen2主干模型进行LoRA微调,RetiBridge仍优于所有对比的开源70亿和320亿模型,在定量准确性、证据依从性、覆盖完整性和BERTScore上表现最佳,并超越OpenAI o3在关键生物标志物相关指标上的表现。代码与数据已开源。
原文摘要 · Abstract (English)
Retinal biomarkers captured by color fundus photography and optical coherence tomography provide clinically valuable evidence for both ocular and systemic diseases. Multimodal large language models (MLLMs) have shown promise for retinal image interpretation, yet existing ophthalmic models rarely quantify these clinically relevant biomarkers or explicitly translate their measurements into qualitative, evidence-grounded diagnostic conclusions. To address this gap, we introduce RetiBridge, a knowledge-guided multimodal large language model that jointly analyzes color fundus photography (CFP), optical coherence tomography (OCT), and text, explicitly bridging quantitative retinal biomarkers to qualitative clinical sub-inferences and coherent diagnostic conclusions. RetiBridge combines knowledge-guided instruction generation, OCT-biomarker alignment, and supervised multimodal instruction tuning to learn a biomarker-grounded quantitative-to-qualitative diagnostic pathway. Using 15,611 paired CFP-OCT samples from UK Biobank with 31 OCT and 6 CFP biomarkers, we construct the Grounded Ophthalmic Understanding benchmark to evaluate diagnostic classification, report generation quality, and fine-grained clinical quality. Despite using only LoRA-based fine-tuning of a 7B-parameter Qwen2 backbone, RetiBridge outperforms all evaluated open-source 7B and 32B baselines, achieving the highest quantitative accuracy, evidence grounding, coverage completeness, and BERTScore, while surpassing OpenAI o3 on these key biomarker-grounded metrics. Our code and data are released in the RetiBridge repository.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。