医学影像中定位关系识别能力不足,现有模型普遍表现不佳。
Your other Left! Vision-Language Models Fail to Identify Relative Positions in Medical Images
- 用标记提示增强视觉输入,但提升有限
- 模型依赖解剖先验而非图像内容,易出错
- 新基准MIRP助力医学定位研究
临床决策高度依赖对解剖结构与病灶相对位置的理解。因此,视觉语言模型(VLMs)要应用于临床,准确判断医学图像中的相对位置是基本前提。尽管如此,这一能力仍严重缺乏研究。我们评估了GPT-4o、Llama3.2、Pixtral和JanusPro等先进VLMs,发现所有模型均无法完成该基础任务。受计算机视觉成功方法启发,我们探索在解剖结构上添加字母数字或彩色标记是否能提升性能。虽然标记带来适度改善,但在医学图像上的表现仍显著低于自然图像。评估表明,医学影像中VLMs更依赖解剖先验知识而非实际图像内容,常导致错误结论。为推动该领域研究,我们提出MIRP(Medical Imaging Relative Positioning)基准数据集,系统评估医学图像中相对位置识别能力。
原文摘要 · Abstract (English)
Clinical decision-making relies heavily on understanding relative positions of anatomical structures and anomalies. Therefore, for Vision-Language Models (VLMs) to be applicable in clinical practice, the ability to accurately determine relative positions on medical images is a fundamental prerequisite. Despite its importance, this capability remains highly underexplored. To address this gap, we evaluate the ability of state-of-the-art VLMs, GPT-4o, Llama3.2, Pixtral, and JanusPro, and find that all models fail at this fundamental task. Inspired by successful approaches in computer vision, we investigate whether visual prompts, such as alphanumeric or colored markers placed on anatomical structures, can enhance performance. While these markers provide moderate improvements, results remain significantly lower on medical images compared to observations made on natural images. Our evaluations suggest that, in medical imaging, VLMs rely more on prior anatomical knowledge than on actual image content for answering relative position questions, often leading to incorrect conclusions. To facilitate further research in this area, we introduce the MIRP , Medical Imaging Relative Positioning, benchmark dataset, designed to systematically evaluate the capability to identify relative positions in medical images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。