arXiv:2503.07094cs.CL2025-03被引 5

构建眼科多模态模型新基准,评估眼底照片与OCT图像诊断能力。

A Novel Ophthalmic Benchmark for Evaluating Multimodal Large Language Models with Fundus Photographs and OCT Images

  • 基于439张眼底图和75张OCT图,建立高质量标注数据集。
  • 7个主流多模态模型在糖尿病视网膜病变上表现较好,但对脉络膜新生血管等病种识别差。
  • 提出临床相关评估框架,助力提升AI在眼科诊疗中的实用价值。

近年来,大语言模型(LLMs)在医疗领域展现出巨大潜力。在此基础上,多模态大语言模型(MLLMs)融合语言与视觉模型,处理包括临床数据和医学影像在内的多种输入。在眼科领域,LLMs已被用于分析光学相干断层扫描(OCT)报告,辅助疾病分类甚至预测治疗结果。然而,现有MLLM基准常无法反映真实临床复杂性,尤其在OCT图像分析方面存在样本量小、数据集多样性不足及专家验证缺失等问题,制约了对模型解读OCT能力的准确评估。我们通过严格质控与专家标注,构建包含439张眼底图像和75张OCT图像的数据集。采用标准化API框架,评估了7个主流MLLMs,发现其在不同疾病上的诊断准确率差异显著:部分模型在糖尿病视网膜病变和年龄相关性黄斑变性中表现良好,但在脉络膜新生血管和近视等疾病上表现不佳,凸显性能不一致问题,亟需进一步优化。研究强调开发临床相关基准的重要性,以更精准评估MLLMs能力。通过持续改进与拓展,有望推动人工智能在眼科诊疗中的实际应用。

原文摘要 · Abstract (English)

In recent years, large language models (LLMs) have demonstrated remarkable potential across various medical applications. Building on this foundation, multimodal large language models (MLLMs) integrate LLMs with visual models to process diverse inputs, including clinical data and medical images. In ophthalmology, LLMs have been explored for analyzing optical coherence tomography (OCT) reports, assisting in disease classification, and even predicting treatment outcomes. However, existing MLLM benchmarks often fail to capture the complexities of real-world clinical practice, particularly in the analysis of OCT images. Many suffer from limitations such as small sample sizes, a lack of diverse OCT datasets, and insufficient expert validation. These shortcomings hinder the accurate assessment of MLLMs' ability to interpret OCT scans and their broader applicability in ophthalmology. Our dataset, curated through rigorous quality control and expert annotation, consists of 439 fundus images and 75 OCT images. Using a standardized API-based framework, we assessed seven mainstream MLLMs and observed significant variability in diagnostic accuracy across different diseases. While some models performed well in diagnosing conditions such as diabetic retinopathy and age-related macular degeneration, they struggled with others, including choroidal neovascularization and myopia, highlighting inconsistencies in performance and the need for further refinement. Our findings emphasize the importance of developing clinically relevant benchmarks to provide a more accurate assessment of MLLMs' capabilities. By refining these models and expanding their scope, we can enhance their potential to transform ophthalmic diagnosis and treatment.

眼科AI多模态模型OCT分析医学评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。