用二选一测试评估医学多模态模型是否真懂图像与文本关联。
Medical-Checklist: Assessing the Comprehension of Medical Images by Multimodal Models

- 给图像配一对一正一误的描述,让模型选对的。
- 四款顶尖模型在常见任务上表现好,但理解能力普遍不足。
- 适合评估模型真实理解力,尤其适用于临床部署前验证。
本文提出新的基准测试 Medical-Checklist,用于评估医学多模态模型的图像理解能力。尽管多模态模型在医学视觉语言任务中展现出巨大潜力,但其性能评估仍具挑战性。核心问题在于:模型能否准确理解输入图像并正确关联相关文本?Medical-Checklist 采用二元判断任务——给定一张图像和两个描述(一个正确,一个错误),模型需选出正确项。错误描述仅替换一个医学概念(词或短语)。该设计虽简单,却能统一评估不同原理训练的模型,并检验其对各类医学概念的跨子领域理解能力。该基准还减少了数据偏差,支持对分布外输入的评估,这是现有数据集难以实现的。对四个领先医学多模态模型的测试发现,尽管它们在 Med-VQA 等特定任务中表现优异,但在理解图像方面仍存在明显缺陷,表明距离临床应用仍有很长路要走。数据集与代码将在论文录用后公开。
原文摘要 · Abstract (English)
This paper introduces a new benchmark test, Medical-Checklist, for assessing medical multimodal models. The recent advancements in multimodal models have demonstrated significant potential in the field of medical vision-language tasks. However, it is becoming increasingly clear that evaluating these models' performance, whether they are applied to natural or medical images, is challenging. The critical question is whether the models can accurately understand an input image while associating it with relevant input text. To address this, Medical-Checklist imposes a binary test on the models: they are given an image and two captions, where one is correct and the other incorrect, and the model must select the correct one. The incorrect caption contains a single medical concept (word or phrase) that is inaccurately substituted from the correct caption. Although the task is simple, this simplicity enables the unified assessment of diverse multimodal models designed and learned on different principles. It also enables us to verify whether models correctly understand a wide range of medical concepts across various medical sub-domains. Medical-Checklist is designed to reduce potential biases in data and to enable evaluation of the models' ability to handle out-of-distribution inputs, which were difficult in existing datasets. When evaluating four state-of-the-art medical multimodal models with Medical-Checklist, it was revealed that despite their excellent performance in specific tasks such as Med-VQA, they may not correctly understand images, suggesting a long journey ahead for clinical application. The dataset and code will be made public upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。