首个面向中文眼科的LLM评测基准,覆盖五大临床场景。
OphthBench: A Comprehensive Benchmark for Evaluating Large Language Models in Chinese Ophthalmology
- 构建五大眼科场景任务,涵盖教育、分诊、诊断等9个具体问题。
- 测试39个主流LLM,发现其在真实临床应用中仍有明显差距。
- 适合医疗AI研究者与眼科临床智能化推进者参考。
大型语言模型(LLMs)在医学领域展现出巨大潜力,眼科是其中重点关注方向。尽管诸多眼科任务已因整合LLMs而显著提升,但在广泛应用于临床前,评估其能力并识别局限性至关重要。为填补这一空白并支持实际应用,我们推出OphthBench——一个专为中文眼科实践设计的综合性评测基准。该基准将典型眼科临床流程划分为五个关键场景:教育、分诊、诊断、治疗和预后。每个场景下设计多种题型,共形成9项任务、591道题目。通过此框架,可全面评估LLMs性能,并揭示其在真实场景中的应用潜力。我们对39个主流LLM进行了系统实验与分析,结果表明当前模型发展与临床实用之间仍存在明显鸿沟,为未来优化指明方向。本工作旨在缩小这一差距,推动眼科领域LLM的持续演进。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown significant promise across various medical applications, with ophthalmology being a notable area of focus. Many ophthalmic tasks have shown substantial improvement through the integration of LLMs. However, before these models can be widely adopted in clinical practice, evaluating their capabilities and identifying their limitations is crucial. To address this research gap and support the real-world application of LLMs, we introduce the OphthBench, a specialized benchmark designed to assess LLM performance within the context of Chinese ophthalmic practices. This benchmark systematically divides a typical ophthalmic clinical workflow into five key scenarios: Education, Triage, Diagnosis, Treatment, and Prognosis. For each scenario, we developed multiple tasks featuring diverse question types, resulting in a comprehensive benchmark comprising 9 tasks and 591 questions. This comprehensive framework allows for a thorough assessment of LLMs' capabilities and provides insights into their practical application in Chinese ophthalmology. Using this benchmark, we conducted extensive experiments and analyzed the results from 39 popular LLMs. Our evaluation highlights the current gap between LLM development and its practical utility in clinical settings, providing a clear direction for future advancements. By bridging this gap, we aim to unlock the potential of LLMs and advance their development in ophthalmology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。