专为眼科设计的开源大模型,临床表现逼近专家水平。
LEME: Open Large Language Models for Ophthalmology with Advanced Reasoning and Clinical Validation
- 两阶段训练:先用20万份临床资料指令微调,再用3万条偏好数据强化学习。
- 在5个零样本任务中超越7个基线模型,比GPT-4o高3.32%(ROUGE-L)。
- 临床评审显示其回答完整度超专家,适用于真实医疗场景辅助决策。
眼病发病率上升带来重大公共卫生负担。大语言模型有望减轻文书工作、支持临床决策,但多数未针对眼科优化,且评估多局限于知识问答,缺乏临床相关基准与真实世界验证。本文提出LEME,一套通过两阶段流程开发的开源大语言模型:(1) 在20万份临床指南、教科书和病例报告上进行指令微调,以增强推理与任务执行能力;(2) 使用约3万条偏好标签进行强化学习,提升准确性与信息量。LEME在五个精心构建的零样本基准上评估,涵盖患者问答、会诊与治疗规划等任务,优于所有七个基线(均p < 0.004),ROUGE-L绝对得分超过GPT-4o 3.32%。进一步使用脱敏患者数据在三个下游任务中评估,由临床医生评审。在患者问答任务中,4项评分(事实性、特异性、完整性、安全性)均获最高分,分别为4.67、4.77、4.79、4.88(1–5分制);完整性得分高于专家撰写答案(4.79 vs. 4.56;p = 0.015)。在视力提取任务中,F1值领先,较LLaMA-3提升14.1%,较Eye-LLaMA提升59.0%。在糖尿病视网膜病变、年龄相关性黄斑变性与青光眼的评估与治疗规划试点中,得分分别为4.36、4.55、4.42、4.36,接近主治医师水平。所有模型、数据与代码将公开,助力后续研发与临床转化,推动诊疗效率与患者照护提升。
原文摘要 · Abstract (English)
The rising prevalence of eye diseases poses a growing public health burden. Large language models (LLMs) offer a promising path to reduce documentation workload and support clinical decision-making. However, few have been tailored for ophthalmology, and most evaluations focus mainly on knowledge-based QA without clinically relevant benchmarks or real-world validation. Here, we present LEME, a suite of open-weight LLMs developed through a two-stage process: (1) instruction tuning on 200,000 samples from clinical guidelines, textbooks, and case reports to enhance reasoning and task-following, and (2) reinforcement learning with ~30,000 preference labels to enhance accuracy and informativeness. LEME was evaluated on five curated zero-shot benchmarks spanning tasks such as patient QA, consultation, and treatment planning. It outperformed all seven baselines (all p < 0.004), exceeding GPT-4o by 3.32% (absolute ROUGE-L gain). It was further evaluated on three downstream tasks using deidentified patient data, reviewed by clinicians. In patient QA, LEME received the highest ratings from attending clinicians in 3 out of 4 criteria, with scores of 4.67 for factuality, 4.77 for specificity, 4.79 for completeness, and 4.88 for safety (1-5 scale). Its completeness score surpassed that of expert-written answers (4.79 vs. 4.56; p = 0.015). In visual acuity extraction, LEME achieved the highest F1, outperforming LLaMA-3 by 14.1% and Eye-LLaMA by 59.0%. In a pilot evaluation on assessment and treatment planning for diabetic retinopathy, AMD, and glaucoma, LEME received scores of 4.36 for factuality, 4.55 for specificity, 4.42 for completeness, and 4.36 for safety, approaching attending-level performance. All models, data, and code will be released to support further development and clinical translation, laying the groundwork for improved efficiency and patient care
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。