测试大模型在意大利医学图像问答中是否真能看懂图片,发现表现差异巨大。
Are Large Vision Language Models Truly Grounded in Medical Images? Evidence from Italian Clinical Visual Question Answering
- 用空白图替代真实医学图像,检验模型是否依赖视觉信息
- GPT-4o准确率下降27.9个百分点,其他模型降幅小至2.4个百分点
- 所有模型都给出看似合理的错误解释,说明存在文本捷径依赖
大型视觉语言模型(VLMs)在医学视觉问答基准上表现优异,但其对视觉信息的依赖程度尚不明确。我们通过欧洲医学问答意大利语数据集中的60个需图像解读的问题,测试四种前沿模型(Claude Sonnet 4.5、GPT-4o、GPT-5-mini、Gemini 2.0 flash exp)的真实视觉理解能力。将正确医学图像替换为空白占位符后,结果显示:GPT-4o视觉依赖最强,准确率下降27.9个百分点(从83.2%降至55.3%);而GPT-5-mini、Gemini和Claude仅下降8.5、2.4和5.6个百分点。分析模型推理过程发现,各模型均能生成看似合理的虚构视觉解释,表明其不同程度依赖文本捷径而非真实视觉分析。研究揭示模型鲁棒性差异,强调临床部署前需严格评估。
原文摘要 · Abstract (English)
Large vision language models (VLMs) have achieved impressive performance on medical visual question answering benchmarks, yet their reliance on visual information remains unclear. We investigate whether frontier VLMs demonstrate genuine visual grounding when answering Italian medical questions by testing four state-of-the-art models: Claude Sonnet 4.5, GPT-4o, GPT-5-mini, and Gemini 2.0 flash exp. Using 60 questions from the EuropeMedQA Italian dataset that explicitly require image interpretation, we substitute correct medical images with blank placeholders to test whether models truly integrate visual and textual information. Our results reveal striking variability in visual dependency: GPT-4o shows the strongest visual grounding with a 27.9pp accuracy drop (83.2% [74.6%, 91.7%] to 55.3% [44.1%, 66.6%]), while GPT-5-mini, Gemini, and Claude maintain high accuracy with modest drops of 8.5pp, 2.4pp, and 5.6pp respectively. Analysis of model-generated reasoning reveals confident explanations for fabricated visual interpretations across all models, suggesting varying degrees of reliance on textual shortcuts versus genuine visual analysis. These findings highlight critical differences in model robustness and the need for rigorous evaluation before clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。