评测4个眼科小模型回答患者问题能力,验证大模型可替代医生评分。
Clinical Validation of Medical-based Large Language Model Chatbots on Ophthalmic Patient Queries with LLM-based Evaluation
- 用眼科患者提问测试4个参数<10亿的小模型,生成2160条回答。
- 最佳模型得分3.44~4.18,但25.5%回答存在幻觉或误导性内容。
- 大模型评分与医生高度一致,适合大规模评估,需人机协同部署。
领域特定的大语言模型在眼科患者教育、分诊和临床决策中日益应用,严格评估对保障安全性和准确性至关重要。本研究评估了四个参数量小于100亿的微型医学LLM(Meerkat-7B、BioMistral-7B、OpenBioLLM-8B、MedLLaMA3-v20)在回答眼科患者问题上的表现,并检验了基于LLM的评估方法相对于临床医生评分的可行性。在横断面研究中,每个模型回答180个眼科患者问题,共生成2160条响应。由三位不同资历的眼科医生及GPT-4-Turbo使用S.C.O.R.E.框架(评估安全性、共识性、客观性、可重复性、可解释性)进行评分,采用五级李克特量表。通过斯皮尔曼等级相关、肯德尔τ统计和核密度估计分析模型与医生评分的一致性。Meerkat-7B表现最优,资深顾问、主治医师和住院医师评分分别为3.44、4.08和4.18。MedLLaMA3-v20表现最差,25.5%的响应包含幻觉或临床误导内容,包括虚构术语。GPT-4-Turbo评分整体与临床医生高度一致,斯皮尔曼ρ=0.80,肯德尔τ=0.67,但资深顾问评分更保守。总体表明医学LLM在眼科问答中具潜在应用价值,但在临床深度和共识性上仍有差距,支持基于LLM的评估可用于大规模基准测试,建议建立混合自动化与临床评审框架以确保安全部署。
原文摘要 · Abstract (English)
Domain specific large language models are increasingly used to support patient education, triage, and clinical decision making in ophthalmology, making rigorous evaluation essential to ensure safety and accuracy. This study evaluated four small medical LLMs Meerkat-7B, BioMistral-7B, OpenBioLLM-8B, and MedLLaMA3-v20 in answering ophthalmology related patient queries and assessed the feasibility of LLM based evaluation against clinician grading. In this cross sectional study, 180 ophthalmology patient queries were answered by each model, generating 2160 responses. Models were selected for parameter sizes under 10 billion to enable resource efficient deployment. Responses were evaluated by three ophthalmologists of differing seniority and by GPT-4-Turbo using the S.C.O.R.E. framework assessing safety, consensus and context, objectivity, reproducibility, and explainability, with ratings assigned on a five point Likert scale. Agreement between LLM and clinician grading was assessed using Spearman rank correlation, Kendall tau statistics, and kernel density estimate analyses. Meerkat-7B achieved the highest performance with mean scores of 3.44 from Senior Consultants, 4.08 from Consultants, and 4.18 from Residents. MedLLaMA3-v20 performed poorest, with 25.5 percent of responses containing hallucinations or clinically misleading content, including fabricated terminology. GPT-4-Turbo grading showed strong alignment with clinician assessments overall, with Spearman rho of 0.80 and Kendall tau of 0.67, though Senior Consultants graded more conservatively. Overall, medical LLMs demonstrated potential for safe ophthalmic question answering, but gaps remained in clinical depth and consensus, supporting the feasibility of LLM based evaluation for large scale benchmarking and the need for hybrid automated and clinician review frameworks to guide safe clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。