arXiv:2503.00901cs.CV2025-03被引 8

构建眼科眼底图像理解评测基准,揭示大模型在识图上的短板

FunBench: Benchmarking Fundus Reading Skills of MLLMs

  • 分四层任务评估眼底图像理解能力,拆解视觉编码与语言模型作用
  • 九个开源模型+GPT-4o测试显示,左右侧识别等基础任务准确率不足
  • 适合医学AI研究者、眼底病诊断系统开发者参考

多模态大语言模型(MLLMs)在医学影像分析中展现出巨大潜力,但在眼底图像解读这一眼科关键技能上仍缺乏有效评估。现有基准未细化任务划分,也未能对大语言模型(LLM)和视觉编码器(VE)进行模块化分析。本文提出FunBench,一个面向眼底图像的视觉问答(VQA)评测基准,采用四层次任务结构(模态感知、解剖感知、病灶分析、疾病诊断),并提供三种评估模式:基于线性探测的VE评估、知识提示的LLM评估、整体综合评估。在九个开源MLLM及GPT-4o上的实验表明,当前模型在基本任务如侧别识别上存在显著缺陷。结果凸显了现有MLLM的局限性,强调需开展领域专用训练以提升模型性能。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have shown significant potential in medical image analysis. However, their capabilities in interpreting fundus images, a critical skill for ophthalmology, remain under-evaluated. Existing benchmarks lack fine-grained task divisions and fail to provide modular analysis of its two key modules, i.e., large language model (LLM) and vision encoder (VE). This paper introduces FunBench, a novel visual question answering (VQA) benchmark designed to comprehensively evaluate MLLMs' fundus reading skills. FunBench features a hierarchical task organization across four levels (modality perception, anatomy perception, lesion analysis, and disease diagnosis). It also offers three targeted evaluation modes: linear-probe based VE evaluation, knowledge-prompted LLM evaluation, and holistic evaluation. Experiments on nine open-source MLLMs plus GPT-4o reveal significant deficiencies in fundus reading skills, particularly in basic tasks such as laterality recognition. The results highlight the limitations of current MLLMs and emphasize the need for domain-specific training and improved LLMs and VEs.

多模态模型医学图像眼底图像评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。