arXiv:2509.25620cs.CV2025-09被引 2

构建眼科多模态大模型评测数据集,助力精准诊疗人工智能发展

LMOD+: A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology

  • 构建含32,633个样本的多模态眼科数据集,覆盖12类常见眼病与5种成像方式
  • 零样本下疾病筛查最高准确率达58%,但疾病分期等任务表现仍不足
  • 支持结构识别、诊断分级与偏见评估,适合医学AI研究者与临床开发者

威胁视力的眼部疾病构成全球重大健康负担,及时诊断受限于专业人才短缺和医疗资源匮乏。尽管多模态大语言模型(MLLMs)在医学图像解析中展现潜力,但其在眼科领域的进展受限于缺乏适合作为生成模型评测基准的综合性数据集。我们提出一个大规模多模态眼科基准,包含32,633个实例,涵盖12种常见眼病及5种影像模态,提供多粒度标注。数据集整合影像、解剖结构、人口统计信息与自由文本注释,支持解剖结构识别、疾病筛查、疾病分期与人口统计预测(用于偏差评估)。本工作在前期LMOD基础上实现三大升级:(1)数据集规模近50%扩展,显著增加彩色眼底照片数量;(2)任务覆盖范围扩大,包括二分类诊断、多分类诊断、按国际标准进行的严重程度分级以及人口统计预测;(3)系统评估24个先进MLLMs。评估结果显示,表现最佳模型在零样本条件下疾病筛查准确率约58%,但在疾病分期等挑战性任务上仍不理想。我们将公开数据集、标注流程与排行榜,以推动眼科AI应用发展,减轻视觉威胁疾病带来的全球负担。

原文摘要 · Abstract (English)

Vision-threatening eye diseases pose a major global health burden, with timely diagnosis limited by workforce shortages and restricted access to specialized care. While multimodal large language models (MLLMs) show promise for medical image interpretation, advancing MLLMs for ophthalmology is hindered by the lack of comprehensive benchmark datasets suitable for evaluating generative models. We present a large-scale multimodal ophthalmology benchmark comprising 32,633 instances with multi-granular annotations across 12 common ophthalmic conditions and 5 imaging modalities. The dataset integrates imaging, anatomical structures, demographics, and free-text annotations, supporting anatomical structure recognition, disease screening, disease staging, and demographic prediction for bias evaluation. This work extends our preliminary LMOD benchmark with three major enhancements: (1) nearly 50% dataset expansion with substantial enlargement of color fundus photography; (2) broadened task coverage including binary disease diagnosis, multi-class diagnosis, severity classification with international grading standards, and demographic prediction; and (3) systematic evaluation of 24 state-of-the-art MLLMs. Our evaluations reveal both promise and limitations. Top-performing models achieved ~58% accuracy in disease screening under zero-shot settings, and performance remained suboptimal for challenging tasks like disease staging. We will publicly release the dataset, curation pipeline, and leaderboard to potentially advance ophthalmic AI applications and reduce the global burden of vision-threatening diseases.

多模态眼科AI大模型评测医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。