arXiv:2603.23953cs.CVcs.ET2026-03被引 2

专为眼科设计的开源多模态大模型,提升疾病诊断与临床决策能力。

VOLMO: Versatile and Open Large Models for Ophthalmology

  • 构建三阶段框架:知识预训练、任务微调、多步临床推理。
  • 在12种眼病上平均F1达87.4%,外部验证表现更优。
  • 模型轻量(20亿参数),开源可用,适合医疗AI研发者使用。

视力损伤影响全球数百万人群,早期检测对防止不可逆失明至关重要。眼科诊疗需整合医学影像、结构化数据与自由文本,过程耗时且负担重。现有通用及医学多模态大模型在眼科表现不佳,且缺乏开放的眼科专用模型。本文提出VOLMO(Versatile and Open Large Models for Ophthalmology),一个模型无关、数据开放的框架,用于构建眼科专用多模态大模型。VOLMO包含三个阶段:在来自82本期刊、26,569篇文献的86,965张图像-文本对上进行眼科知识预训练;在涵盖12种眼病的26,929个标注样本上进行领域任务微调,实现疾病筛查与分期分类;在913份患者病例报告上进行多步临床推理,完成评估、规划与随访。基于该框架,我们训练了一个20亿参数的紧凑型多模态大模型,并与InternVL-2B、LLaVA-Med-7B、MedGemma-4B、MedGemma-27B、RETFound等强基线对比。评估涵盖图像描述生成、疾病筛查与分期分类、评估与管理生成,另经两名医疗专业人员人工评审,并在三个独立队列中对老年性黄斑变性和糖尿病视网膜病变进行外部验证。在各类场景下,VOLMO-2B均持续优于基线,图像描述表现更强,12种眼病平均F1达87.4%,外部验证得分更高。

原文摘要 · Abstract (English)

Vision impairment affects millions globally, and early detection is critical to preventing irreversible vision loss. Ophthalmology workflows require clinicians to integrate medical images, structured clinical data, and free-text notes to determine disease severity and management, which is time-consuming and burdensome. Recent multimodal large language models (MLLMs) show promise, but existing general and medical MLLMs perform poorly in ophthalmology, and few ophthalmology-specific MLLMs are openly available. We present VOLMO (Versatile and Open Large Models for Ophthalmology), a model-agnostic, data-open framework for developing ophthalmology-specific MLLMs. VOLMO includes three stages: ophthalmology knowledge pretraining on 86,965 image-text pairs from 26,569 articles across 82 journals; domain task fine-tuning on 26,929 annotated instances spanning 12 eye conditions for disease screening and severity classification; and multi-step clinical reasoning on 913 patient case reports for assessment, planning, and follow-up care. Using this framework, we trained a compact 2B-parameter MLLM and compared it with strong baselines, including InternVL-2B, LLaVA-Med-7B, MedGemma-4B, MedGemma-27B, and RETFound. We evaluated these models on image description generation, disease screening and staging classification, and assessment-and-management generation, with additional manual review by two healthcare professionals and external validation on three independent cohorts for age-related macular degeneration and diabetic retinopathy. Across settings, VOLMO-2B consistently outperformed baselines, achieving stronger image description performance, an average F1 of 87.4% across 12 eye conditions, and higher scores in external validation.

眼科AI多模态模型临床推理开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。