构建首个全面评估多模态大模型看眼底OCT的基准,揭示现有模型仍远未达临床理解水平。
Can Multimodal Large Language Models Understand OCT?

- 设计分层任务体系,覆盖从图像感知到临床决策的20项细粒度能力
- 在10,076道题、4,137张OCT图像上测试20个主流模型,性能普遍不理想
- 发现医学领域微调和模型规模增大均不能稳定提升理解力,适合研究者与医生参考
光学相干断层扫描(OCT)对视网膜疾病诊断至关重要。尽管多模态大语言模型(MLLMs)在医学图像分析中展现出潜力,但现有评测多局限于粗粒度疾病分类或孤立问答,未能充分评估从视觉感知到临床推理的完整认知过程。为此,我们提出OCT-Bench,一个专注于OCT图像理解的综合性基准。该基准包含从7个公开数据集的4,137张OCT图像中构建的10,076道高质量选择题。依据真实临床解读流程,建立涵盖感知、认知、推理三个维度的20项细粒度任务体系,覆盖成像特征、视网膜解剖、病灶特性、空间关系、疾病评估、治疗决策与预后管理等。我们系统评估了20个代表性MLLMs,包括专有模型、开源通用模型及医疗领域模型。实验表明,当前模型在可靠OCT理解方面仍有显著差距;且医学领域适配或模型规模扩大,并未在各能力层级上一致提升表现。OCT-Bench为全面、精细地评估MLLMs提供了基础,有助于识别能力瓶颈,推动面向临床的OCT理解发展。
原文摘要 · Abstract (English)
Optical coherence tomography (OCT) imaging is essential for the diagnosis and treatment of retinal diseases. Although multimodal large language models (MLLMs) have demonstrated considerable potential in medical image analysis, existing benchmarks largely reduce OCT understanding to coarse-grained disease classification or isolated visual question answering, leaving the complete cognitive process from visual perception to clinical reasoning insufficiently evaluated. To address this limitation, we introduce OCT-Bench, a comprehensive benchmark dedicated to OCT image understanding. OCT-Bench comprises 10,076 high-quality multiple-choice questions constructed from 4,137 OCT images across seven public datasets. Following the real-world clinical interpretation workflow, we establish a hierarchical capability taxonomy consisting of 20 fine-grained tasks across three dimensions: Perception, Cognition, and Reasoning. These tasks cover a broad range of capabilities, including imaging attributes, retinal anatomy, lesion characteristics, spatial relationships, disease assessment, therapeutic decision-making, and prognostic management. We systematically evaluate 20 representative MLLMs, including proprietary models, open-source general-purpose models, and medical-domain models. Experimental results demonstrate that current models remain substantially short of reliable OCT understanding. Moreover, neither medical-domain adaptation nor increased model scale consistently improves performance across capability levels. OCT-Bench enables comprehensive and fine-grained evaluation of MLLMs, providing a foundation for identifying capability bottlenecks and advancing clinically grounded OCT understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。