首个全面评估内镜多模态大模型的基准,覆盖真实临床全流程。
EndoBench: A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy Analysis
- 构建4类内镜场景+12项任务,含6832个严格验证的图文问答对。
- 专有模型表现最佳但仍不及人类专家,微调显著提升准确率。
- 适合医学AI研究者、临床工程师及模型开发者参考使用。
内镜检查是诊断和治疗内部疾病的关键手段,多模态大语言模型(MLLMs)正日益用于辅助内镜分析。然而,现有基准普遍局限在特定内镜场景和少量临床任务,无法反映真实世界中内镜实践的多样性与临床工作流所需的完整能力。为此,我们提出EndoBench,首个专门设计用于全面评估MLLMs在内镜分析全谱系表现的综合性基准。EndoBench涵盖4种不同内镜场景、12项专业临床任务(含12个子任务)以及5级视觉提示粒度,共包含6,832个经过严格验证的视觉问答对,数据源自21个多样化数据集。其多维度评估框架模拟真实临床流程——从解剖识别、病灶分析、空间定位到手术操作——全面衡量MLLM在真实场景中的感知与诊断能力。我们对23个顶尖模型(包括通用型、医疗专用型及专有模型)进行评测,并以人类临床医生表现作为参照标准。实验结果表明:(1) 专有模型整体优于开源与医疗专用模型,但仍落后于人类专家;(2) 医疗领域监督微调显著提升任务特定准确性;(3) 模型性能对提示格式与任务复杂度仍敏感。EndoBench为推进内镜领域MLLM的评估与发展树立了新标准,揭示了当前模型与专家推理之间的进步与差距。我们已公开发布该基准及代码。
原文摘要 · Abstract (English)
Endoscopic procedures are essential for diagnosing and treating internal diseases, and multi-modal large language models (MLLMs) are increasingly applied to assist in endoscopy analysis. However, current benchmarks are limited, as they typically cover specific endoscopic scenarios and a small set of clinical tasks, failing to capture the real-world diversity of endoscopic scenarios and the full range of skills needed in clinical workflows. To address these issues, we introduce EndoBench, the first comprehensive benchmark specifically designed to assess MLLMs across the full spectrum of endoscopic practice with multi-dimensional capacities. EndoBench encompasses 4 distinct endoscopic scenarios, 12 specialized clinical tasks with 12 secondary subtasks, and 5 levels of visual prompting granularities, resulting in 6,832 rigorously validated VQA pairs from 21 diverse datasets. Our multi-dimensional evaluation framework mirrors the clinical workflow--spanning anatomical recognition, lesion analysis, spatial localization, and surgical operations--to holistically gauge the perceptual and diagnostic abilities of MLLMs in realistic scenarios. We benchmark 23 state-of-the-art models, including general-purpose, medical-specialized, and proprietary MLLMs, and establish human clinician performance as a reference standard. Our extensive experiments reveal: (1) proprietary MLLMs outperform open-source and medical-specialized models overall, but still trail human experts; (2) medical-domain supervised fine-tuning substantially boosts task-specific accuracy; and (3) model performance remains sensitive to prompt format and clinical task complexity. EndoBench establishes a new standard for evaluating and advancing MLLMs in endoscopy, highlighting both progress and persistent gaps between current models and expert clinical reasoning. We publicly release our benchmark and code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。