构建眼科多模态数据集LMOD,评测大模型在眼病诊断中的表现
LMOD: A Large Multimodal Ophthalmology Dataset and Benchmark for Large Vision-Language Models
- 构建包含2.2万例的多模态眼科数据集,覆盖五类影像与文本信息
- 大模型在眼科任务中性能显著下降,平均准确率远低于专用神经网络
- 揭示六类关键错误模式,推动面向眼科的大视觉语言模型研发
眼科致盲性疾病在全球范围内负担沉重,许多病例未能及时诊断。大视觉语言模型(LVLMs)有望辅助理解解剖结构、诊断眼病并生成诊疗建议,减轻临床压力并提升眼健康可及性。然而,现有评估基准严重不足。本研究提出LMOD,一个大规模多模态眼科基准,包含21,993个实例,涵盖五种眼科成像模态(光学相干断层扫描、彩色眼底照片、扫描激光眼底镜、晶状体照片、术中影像)、自由文本、人口统计学信息及疾病生物标志物,并覆盖解剖理解、疾病诊断与亚组分析等核心应用。我们对13个来自闭源、开源及医疗领域的先进LVLM进行了评测。结果显示,相比其他领域,LVLM在眼科任务中性能显著下降。系统性错误分析识别出六类主要失败模式:误分类、拒绝回答失败、推理不一致、幻觉、无依据断言及缺乏领域知识。相比之下,针对这些任务专门训练的监督神经网络表现出高准确率。这些发现凸显了开发眼科专用基准的紧迫性。
原文摘要 · Abstract (English)
The prevalence of vision-threatening eye diseases is a significant global burden, with many cases remaining undiagnosed or diagnosed too late for effective treatment. Large vision-language models (LVLMs) have the potential to assist in understanding anatomical information, diagnosing eye diseases, and drafting interpretations and follow-up plans, thereby reducing the burden on clinicians and improving access to eye care. However, limited benchmarks are available to assess LVLMs' performance in ophthalmology-specific applications. In this study, we introduce LMOD, a large-scale multimodal ophthalmology benchmark consisting of 21,993 instances across (1) five ophthalmic imaging modalities: optical coherence tomography, color fundus photographs, scanning laser ophthalmoscopy, lens photographs, and surgical scenes; (2) free-text, demographic, and disease biomarker information; and (3) primary ophthalmology-specific applications such as anatomical information understanding, disease diagnosis, and subgroup analysis. In addition, we benchmarked 13 state-of-the-art LVLM representatives from closed-source, open-source, and medical domains. The results demonstrate a significant performance drop for LVLMs in ophthalmology compared to other domains. Systematic error analysis further identified six major failure modes: misclassification, failure to abstain, inconsistent reasoning, hallucination, assertions without justification, and lack of domain-specific knowledge. In contrast, supervised neural networks specifically trained on these tasks as baselines demonstrated high accuracy. These findings underscore the pressing need for benchmarks in the development and validation of ophthalmology-specific LVLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。