arXiv:2603.14323cs.CVcs.AI2026-03被引 3

医学多模态模型常错把图像区域,研究发现并提出改进方法。

How Do Medical MLLMs Fail? A Study on Visual Grounding in Medical Images

  • 构建专家指导的VGMED数据集,专门评测模型对医学图像的视觉定位能力。
  • 8个顶尖模型在110万+样本上均未能准确定位病灶区域。
  • 无需训练的推理优化方法VGRefine可显著提升医学图像理解性能。

通用多模态大模型在跨视觉-语言任务中表现优异,但在医疗任务尤其是零样本场景下表现不佳。关键原因是缺乏对医学多模态大模型(MLLMs)在医学图像理解中失败原因的系统性理解。本文首次系统研究了前沿医学MLLMs的视觉定位能力,设计了由临床专家指导的VGMED数据集,专门评估其在医学图像中的视觉接地能力,并引入新量化指标与定性分析。实验覆盖8个最先进的医学MLLMs,在超过11万张来自8种成像模态的医学多模态问答样本上验证:这些模型常无法将预测锚定在临床相关图像区域。这一现象在自然图像中并不明显,说明其问题具有医学特异性。基于此,我们提出一种简单有效的推理阶段方法VGRefine,通过优化注意力分布来增强视觉定位能力。该方法在6个不同医学视觉问答基准上达到当前最优表现,且无需额外训练或外部专家模型。本工作首次系统证实,视觉定位不足是医学MLLMs性能欠佳的关键因素之一。

原文摘要 · Abstract (English)

Generalist multimodal large language models (MLLMs) have achieved impressive performance across a wide range of vision-language tasks. However, their performance on medical tasks, particularly in zero-shot settings where generalization is critical, remains suboptimal. A key research gap is the limited understanding of why medical MLLMs underperform in medical image interpretation. In this work, we present a pioneering systematic investigation into the visual grounding capabilities of state-of-the-art medical MLLMs. To disentangle visual grounding from semantic grounding, we design VGMED, a novel evaluation dataset developed with expert clinical guidance, explicitly assessing the visual grounding capability of medical MLLMs. We introduce new quantitative metrics and conduct detailed qualitative analyses. Our study across eight state-of-the-art (SOTA) medical MLLMs validates that they often fail to ground their predictions in clinically relevant image regions. We note that this finding is specific to medical image analysis; in contrast, prior work has shown that MLLMs are capable of grounding their predictions in the correct image regions when applied to natural scene images. Motivated by these findings, we propose VGRefine, a simple yet effective inference-time method that refines attention distribution to improve visual grounding in medical settings. Our approach achieves SOTA performance across 6 diverse Med-VQA benchmarks (over 110K VQA samples from 8 imaging modalities) without requiring additional training or external expert models. Overall, our work, for the first time, systematically validates inadequate visual grounding as one of the key contributing factors for medical MLLMs' under-performance. Additional experiments are included in the Supp.

医学多模态视觉定位模型缺陷分析推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。