开源视觉语言模型在多种医学影像诊断中表现不一,尤其在胸片上表现优异。
Diagnostic Accuracy of Open-Source Vision-Language Models on Diverse Medical Imaging Tasks
- 用多模态输入和思维链提示测试5个开源模型的诊断能力
- Qwen2.5在胸片和内镜图像上准确率达90.4%和84.2%,显著领先
- 模型在眼底图像任务中表现差,且多模态和思维链未提升效果
本回顾性研究使用MedFMC数据集评估了五个视觉语言模型(Qwen2.5、Phi-4、Gemma3、Llama3.2和Mistral3.1)在多种医学影像任务中的诊断准确性。该数据集包含来自7,461名患者的22,349张图像,涵盖胸部放射摄影(19种疾病多标签分类)、结肠病理(肿瘤检测)、内窥镜检查(结直肠病变识别)、新生儿黄疸评估(基于皮肤颜色判断治疗必要性)和视网膜眼底照相(5级糖尿病视网膜病变分级)。在仅视觉输入、多模态输入和思维链推理三种实验设置下,模型准确性与真实标签对比,采用自助法置信区间进行统计检验(p<.05)。Qwen2.5在胸部放射摄影(90.4%)和内窥镜图像(84.2%)上表现最佳,显著优于其他模型(p<.001)。在结肠病理任务中,Qwen2.5(69.0%)与Phi-4(69.6%)表现相近(p=.41),均显著高于其他模型(p<.001);新生儿黄疸评估中,两者(分别为58.3%和58.1%)也表现相当且显著领先(p<.001)。所有模型在视网膜眼底照相任务中表现不佳,其中Qwen2.5和Gemma3最高,为18.6%(相近,p=.99),仍显著优于其他模型(p<.001)。出乎意料的是,多模态输入降低了部分模型和任务的准确性,思维链提示也未能提升性能。开源视觉语言模型在胸部放射摄影等任务中展现出良好诊断潜力,但在复杂领域如视网膜图像分析中仍有限,需进一步开发与领域适配后方可临床应用。
原文摘要 · Abstract (English)
This retrospective study evaluated five VLMs (Qwen2.5, Phi-4, Gemma3, Llama3.2, and Mistral3.1) using the MedFMC dataset. This dataset includes 22,349 images from 7,461 patients encompassing chest radiography (19 disease multi-label classifications), colon pathology (tumor detection), endoscopy (colorectal lesion identification), neonatal jaundice assessment (skin color-based treatment necessity), and retinal fundoscopy (5-point diabetic retinopathy grading). Diagnostic accuracy was compared in three experimental settings: visual input only, multimodal input, and chain-of-thought reasoning. Model accuracy was assessed against ground truth labels, with statistical comparisons using bootstrapped confidence intervals (p<.05). Qwen2.5 achieved the highest accuracy for chest radiographs (90.4%) and endoscopy images (84.2%), significantly outperforming the other models (p<.001). In colon pathology, Qwen2.5 (69.0%) and Phi-4 (69.6%) performed comparably (p=.41), both significantly exceeding other VLMs (p<.001). Similarly, for neonatal jaundice assessment, Qwen2.5 (58.3%) and Phi-4 (58.1%) showed comparable leading accuracies (p=.93) significantly exceeding their counterparts (p<.001). All models struggled with retinal fundoscopy; Qwen2.5 and Gemma3 achieved the highest, albeit modest, accuracies at 18.6% (comparable, p=.99), significantly better than other tested models (p<.001). Unexpectedly, multimodal input reduced accuracy for some models and modalities, and chain-of-thought reasoning prompts also failed to improve accuracy. The open-source VLMs demonstrated promising diagnostic capabilities, particularly in chest radiograph interpretation. However, performance in complex domains such as retinal fundoscopy was limited, underscoring the need for further development and domain-specific adaptation before widespread clinical application.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。