arXiv:2507.11200cs.CV2025-07被引 3

对比了8个医疗视觉语言模型,发现通用大模型已可胜任部分医学任务。

How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study

  • 按理解与推理能力拆解模型表现,发现推理短板明显。
  • 最大模型在多个数据集上超越专用模型,零样本迁移能力出色。
  • 当前无模型适合临床部署,需更强对齐与精细评估。

在自然图像任务中表现出色的视觉语言模型(VLMs)正被引入医疗领域,但其在医学任务中的实际能力尚不明确。本文对开源通用及医疗专用的VLMs(参数量3B至72B)进行了全面评估,涵盖8个基准:MedXpert、OmniMedVQA、PMC-VQA、PathVQA、MMMU、SLAKE和VQA-RAD。通过分离理解与推理能力,发现三方面关键结论:第一,大尺寸通用模型在多个基准上已达或超过专用模型水平,显示自然图像到医学图像的强零样本迁移能力;第二,推理能力始终低于理解能力,构成临床决策支持的主要障碍;第三,各基准表现差异显著,反映任务设计、标注质量与知识需求的不同。目前无模型达到临床部署可靠性标准,亟需更强的多模态对齐与更严谨、细粒度的评估体系。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) trained on web-scale corpora excel at natural image tasks and are increasingly repurposed for healthcare; however, their competence in medical tasks remains underexplored. We present a comprehensive evaluation of open-source general-purpose and medically specialised VLMs, ranging from 3B to 72B parameters, across eight benchmarks: MedXpert, OmniMedVQA, PMC-VQA, PathVQA, MMMU, SLAKE, and VQA-RAD. To observe model performance across different aspects, we first separate it into understanding and reasoning components. Three salient findings emerge. First, large general-purpose models already match or surpass medical-specific counterparts on several benchmarks, demonstrating strong zero-shot transfer from natural to medical images. Second, reasoning performance is consistently lower than understanding, highlighting a critical barrier to safe decision support. Third, performance varies widely across benchmarks, reflecting differences in task design, annotation quality, and knowledge demands. No model yet reaches the reliability threshold for clinical deployment, underscoring the need for stronger multimodal alignment and more rigorous, fine-grained evaluation protocols.

视觉语言模型医疗AI基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。