测试医学多模态模型的基础感知能力,发现其表现远低于人类。
MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine
- 构建涵盖8类临床任务的基准测试MedBLINK,评估模型对图像方向、增强类型等基础感知能力。
- 19个主流模型在1429道题上最高仅达65%准确率,人类为96.4%。
- 揭示当前医学多模态模型视觉理解不足,限制临床应用前景。
多模态语言模型(MLMs)在临床决策支持和诊断推理中展现潜力,有望实现医学影像的端到端自动化解读。然而,临床医生对AI工具极为谨慎,若模型在图像方向判断、是否含对比剂等基础感知任务上出错,将难以获得信任。为此,我们提出MedBLINK,一个用于探测此类感知能力的基准测试。该基准覆盖多种成像模态与解剖区域的8类临床任务,共包含1,429道多项选择题,基于1,605张图像。我们评估了19个顶尖的MLM,包括通用模型(GPT-4o、Claude 3.5 Sonnet)和领域专用模型(Med Flamingo、LLaVA Med、RadFM)。尽管人类标注者达到96.4%的准确率,表现最好的模型也仅达65%。结果表明,当前模型在常规感知任务上频繁失败,亟需加强视觉语义对齐以推动临床采纳。数据可在项目页面获取。
原文摘要 · Abstract (English)
Multimodal language models (MLMs) show promise for clinical decision support and diagnostic reasoning, raising the prospect of end-to-end automated medical image interpretation. However, clinicians are highly selective in adopting AI tools; a model that makes errors on seemingly simple perception tasks such as determining image orientation or identifying whether a CT scan is contrast-enhance are unlikely to be adopted for clinical tasks. We introduce Medblink, a benchmark designed to probe these models for such perceptual abilities. Medblink spans eight clinically meaningful tasks across multiple imaging modalities and anatomical regions, totaling 1,429 multiple-choice questions over 1,605 images. We evaluate 19 state-of-the-art MLMs, including general purpose (GPT4o, Claude 3.5 Sonnet) and domain specific (Med Flamingo, LLaVA Med, RadFM) models. While human annotators achieve 96.4% accuracy, the best-performing model reaches only 65%. These results show that current MLMs frequently fail at routine perceptual checks, suggesting the need to strengthen their visual grounding to support clinical adoption. Data is available on our project page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。