将语言与视觉模型融合,实现眼病自动诊断与分诊
An integrated language-vision foundation model for conversational diagnostics and triaging in primary eye care
- 用路由机制根据文本查询匹配专用视觉模型
- 疾病检测准确率超82.2%,严重程度区分超89%
- 适合基层医疗辅助诊断或在线眼底评估
当前深度学习模型多为特定任务设计,且缺乏友好的操作界面。我们提出Meta-EyeFM,一种集成大型语言模型(LLM)与视觉基础模型(VFMs)的多功能基础模型,用于眼部疾病评估。该模型通过路由机制根据文本查询精准调用对应视觉模型。采用低秩适配(LoRA)对视觉模型进行微调,实现眼部及全身疾病检测、疾病严重程度区分和常见眼部体征识别。模型在将眼底图像路由至相应视觉模型时达到100%准确率,各模型在疾病检测中准确率≥82.2%,严重程度区分≥89%,体征识别≥76%。相较于Gemini-1.5-flash和ChatGPT-4o LMMs,Meta-EyeFM在多种眼病检测上高出11%至43%,性能接近专业眼科医生。该系统提升了易用性与诊断性能,可作为初级眼保健的决策支持工具或在线眼底评估大模型。
原文摘要 · Abstract (English)
Current deep learning models are mostly task specific and lack a user-friendly interface to operate. We present Meta-EyeFM, a multi-function foundation model that integrates a large language model (LLM) with vision foundation models (VFMs) for ocular disease assessment. Meta-EyeFM leverages a routing mechanism to enable accurate task-specific analysis based on text queries. Using Low Rank Adaptation, we fine-tuned our VFMs to detect ocular and systemic diseases, differentiate ocular disease severity, and identify common ocular signs. The model achieved 100% accuracy in routing fundus images to appropriate VFMs, which achieved $\ge$ 82.2% accuracy in disease detection, $\ge$ 89% in severity differentiation, $\ge$ 76% in sign identification. Meta-EyeFM was 11% to 43% more accurate than Gemini-1.5-flash and ChatGPT-4o LMMs in detecting various eye diseases and comparable to an ophthalmologist. This system offers enhanced usability and diagnostic performance, making it a valuable decision support tool for primary eye care or an online LLM for fundus evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。