首个全景式胃肠道内镜多模态大模型评测基准,揭示模型在临床实用中的关键缺陷。
GI-Bench: A Panoramic Benchmark Revealing the Knowledge-Experience Dissociation of Multimodal Large Language Models in Gastrointestinal Endoscopy Against Clinical Standards
- 构建涵盖20类病灶的全景评测框架,覆盖五阶段内镜临床流程。
- 顶尖模型诊断准确率超训练医师(0.641 vs 0.492),但定位能力显著落后(mIoU 0.345 vs 0.506)。
- 模型报告语言流畅但事实错误多,存在‘过度解读’与幻觉问题,适合研究者参考。
多模态大语言模型(MLLMs)在胃肠病学中展现出潜力,但其在完整临床流程和人类标准下的表现尚未验证。为系统评估前沿MLLMs在全景胃肠道内镜流程中的表现,并比较其临床实用性与人类内镜医师的差距,我们构建了GI-Bench基准,包含20个细粒度病灶类别。12个MLLM在五个临床阶段(解剖定位、病灶识别、诊断、发现描述、管理建议)进行评估,性能与三名初级内镜医师及三名住院医师对比,采用宏平均F1、均交并比(mIoU)及多维李克特量表。Gemini-3-Pro表现最佳。在诊断推理中,顶级模型(宏平均F1 0.641)优于住院医师(0.492),接近初级医师(0.727;p>0.05)。然而,存在严重“空间定位瓶颈”:人类定位准确率(mIoU >0.506)显著高于最优模型(0.345;p<0.05)。定性分析显示“流畅性-准确性悖论”:模型报告语言可读性优于人类(p<0.05),但事实正确率更低(p<0.05),因对视觉特征存在“过度解读”与幻觉。GI-Bench维护动态排行榜,实时追踪模型在临床内镜中的演进。当前排名与结果可访问 https://roterdl.github.io/GIBench/。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) show promise in gastroenterology, yet their performance against comprehensive clinical workflows and human benchmarks remains unverified. To systematically evaluate state-of-the-art MLLMs across a panoramic gastrointestinal endoscopy workflow and determine their clinical utility compared with human endoscopists. We constructed GI-Bench, a benchmark encompassing 20 fine-grained lesion categories. Twelve MLLMs were evaluated across a five-stage clinical workflow: anatomical localization, lesion identification, diagnosis, findings description, and management. Model performance was benchmarked against three junior endoscopists and three residency trainees using Macro-F1, mean Intersection-over-Union (mIoU), and multi-dimensional Likert scale. Gemini-3-Pro achieved state-of-the-art performance. In diagnostic reasoning, top-tier models (Macro-F1 0.641) outperformed trainees (0.492) and rivaled junior endoscopists (0.727; p>0.05). However, a critical "spatial grounding bottleneck" persisted; human lesion localization (mIoU >0.506) significantly outperformed the best model (0.345; p<0.05). Furthermore, qualitative analysis revealed a "fluency-accuracy paradox": models generated reports with superior linguistic readability compared with humans (p<0.05) but exhibited significantly lower factual correctness (p<0.05) due to "over-interpretation" and hallucination of visual features. GI-Bench maintains a dynamic leaderboard that tracks the evolving performance of MLLMs in clinical endoscopy. The current rankings and benchmark results are available at https://roterdl.github.io/GIBench/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。