测试医学多模态模型在视觉任务中的可靠性,发现现有模型表现低于随机猜测。
MediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation models
- 构建视觉问答数据集MediConfusion,从视觉角度探测模型失效模式。
- 所有主流模型在该数据集上表现均低于随机水平,准确率不足50%。
- 揭示模型对细微图像差异的误判规律,为可信医疗AI设计提供参考。
多模态大语言模型(MLLMs)有望通过自动化解决方案提升医疗诊断的准确性、可及性和成本效益。尽管近年来医学MLLMs取得进展,但其能力与局限性仍不明确。现有基准多聚焦于通用医学知识,却严重忽视了模型在安全关键领域中的系统性失败模式。本文提出MediConfusion,一个挑战性的医学视觉问答(VQA)基准数据集,从视觉角度探测医学MLLMs的缺陷。结果表明,最先进的模型在视觉差异明显、但对医生清晰可辨的图像对上极易混淆。令人震惊的是,所有现成模型(开源或专有)在MediConfusion上的表现均低于随机猜测,暴露出当前医学MLLMs在医疗部署中的可靠性危机。我们还提取了模型失败的常见模式,为下一代更可信、可靠的医疗AI模型设计提供依据。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have tremendous potential to improve the accuracy, availability, and cost-effectiveness of healthcare by providing automated solutions or serving as aids to medical professionals. Despite promising first steps in developing medical MLLMs in the past few years, their capabilities and limitations are not well-understood. Recently, many benchmark datasets have been proposed that test the general medical knowledge of such models across a variety of medical areas. However, the systematic failure modes and vulnerabilities of such models are severely underexplored with most medical benchmarks failing to expose the shortcomings of existing models in this safety-critical domain. In this paper, we introduce MediConfusion, a challenging medical Visual Question Answering (VQA) benchmark dataset, that probes the failure modes of medical MLLMs from a vision perspective. We reveal that state-of-the-art models are easily confused by image pairs that are otherwise visually dissimilar and clearly distinct for medical experts. Strikingly, all available models (open-source or proprietary) achieve performance below random guessing on MediConfusion, raising serious concerns about the reliability of existing medical MLLMs for healthcare deployment. We also extract common patterns of model failure that may help the design of a new generation of more trustworthy and reliable MLLMs in healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。