arXiv:2505.08414eess.IVcs.CV2025-05被引 6

将语言与视觉模型融合,实现眼病自动诊断与分诊

An integrated language-vision foundation model for conversational diagnostics and triaging in primary eye care

  • 用路由机制根据文本查询匹配专用视觉模型
  • 疾病检测准确率超82.2%,严重程度区分超89%
  • 适合基层医疗辅助诊断或在线眼底评估

当前深度学习模型多为特定任务设计,且缺乏友好的操作界面。我们提出Meta-EyeFM,一种集成大型语言模型(LLM)与视觉基础模型(VFMs)的多功能基础模型,用于眼部疾病评估。该模型通过路由机制根据文本查询精准调用对应视觉模型。采用低秩适配(LoRA)对视觉模型进行微调,实现眼部及全身疾病检测、疾病严重程度区分和常见眼部体征识别。模型在将眼底图像路由至相应视觉模型时达到100%准确率,各模型在疾病检测中准确率≥82.2%,严重程度区分≥89%,体征识别≥76%。相较于Gemini-1.5-flash和ChatGPT-4o LMMs,Meta-EyeFM在多种眼病检测上高出11%至43%,性能接近专业眼科医生。该系统提升了易用性与诊断性能,可作为初级眼保健的决策支持工具或在线眼底评估大模型。

原文摘要 · Abstract (English)

Current deep learning models are mostly task specific and lack a user-friendly interface to operate. We present Meta-EyeFM, a multi-function foundation model that integrates a large language model (LLM) with vision foundation models (VFMs) for ocular disease assessment. Meta-EyeFM leverages a routing mechanism to enable accurate task-specific analysis based on text queries. Using Low Rank Adaptation, we fine-tuned our VFMs to detect ocular and systemic diseases, differentiate ocular disease severity, and identify common ocular signs. The model achieved 100% accuracy in routing fundus images to appropriate VFMs, which achieved $\ge$ 82.2% accuracy in disease detection, $\ge$ 89% in severity differentiation, $\ge$ 76% in sign identification. Meta-EyeFM was 11% to 43% more accurate than Gemini-1.5-flash and ChatGPT-4o LMMs in detecting various eye diseases and comparable to an ophthalmologist. This system offers enhanced usability and diagnostic performance, making it a valuable decision support tool for primary eye care or an online LLM for fundus evaluation.

眼病诊断多模态模型医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。