评测6种视觉语言模型在真实伤口图像上的临床分析能力,发现通用模型表现优于专业模型。
Evaluation of Medical Vision Language Models HuluMed and MedGemma, and general purpose chatbots Gemma 3, ChatGPT Plus, and Claude Pro on real previously unseen wound images
- 用12个临床问题框架评估模型对20例真实伤口的多模态理解与推理能力
- 通用模型ChatGPT准确率达72.5%,显著高于医疗专用模型最高40%的表现
- 尽管有潜力,现有模型在治疗规划和自主决策上仍存在明显局限
慢性伤口评估因需精准解读创面形态、组织成分、血管特征及感染风险而具临床挑战性。近期视觉语言模型(VLMs)的发展使结合图像理解与临床推理的自动化多模态分析成为可能。本研究基于涵盖血管性、外科性、缺血性、静脉性、淋巴水肿及截肢相关等20类临床多样伤口的扩展数据集,评估了六种通用与医学专用开源及专有VLMs的临床伤口评估性能。采用包含12个临床问题的结构化框架,涵盖伤口分类、感染风险、血管干预建议、清创紧迫性、治疗方案选择及高级管理计划等维度。在20个伤口案例、240项临床评分决策中,ChatGPT表现最佳,正确率达174/240(72.50%),其次为Claude(149/240,62.08%)。在开源及医学专用模型中,HuluMed表现最优(96/240,40.00%),其次为Gemma 3(81/240,33.75%)、MedGemma 4B(62/240,25.83%)、MedGemma 27B(42/240,17.50%)。结果表明,当前前沿通用多模态系统在伤口分析上显著优于医疗专用模型,凸显广泛多模态推理能力与领域知识并重的重要性。尽管现有VLMs展现出临床决策支持的潜力,但在高级伤口管理推理、操作规划及自主临床可靠性方面仍存在显著不足。
原文摘要 · Abstract (English)
Chronic wound assessment remains a clinically challenging task that requires accurate interpretation of wound morphology, tissue composition, vascular characteristics, and infection risk. Recent advances in Vision-Language Models (VLMs) have introduced the possibility of automated multimodal wound analysis through image understanding combined with clinical reasoning. This study evaluates the performance of several general-purpose and medically specialized open-source and proprietary VLMs for clinical wound assessment using an expanded, curated dataset of 20 clinically diverse wounds spanning vascular, surgical, ischemic, venous, lymphedema, and amputation-related etiologies. Six VLMs were evaluated using a structured twelve-question clinical framework covering wound classification, infection risk, vascular intervention recommendations, debridement urgency, wound therapy selection, and advanced management planning. Across 20 wound cases and 240 clinician-graded wound-analysis decisions, ChatGPT achieved the highest overall performance with 174/240 correct responses (72.50%), followed by Claude with 149/240 (62.08%). Among the open-source and medically specialized models, HuluMed achieved the strongest performance with 96/240 correct responses (40.00%), followed by Gemma 3 (81/240, 33.75%), MedGemma 4B (62/240, 25.83%), and MedGemma 27B (42/240, 17.50%). The findings suggest that frontier general-purpose multimodal systems currently demonstrate substantially stronger wound-analysis performance than medically specialized alternatives, highlighting the continued importance of broad multimodal reasoning capabilities alongside domain-specific medical knowledge. Although current VLMs demonstrate promising potential for clinical decision support, substantial limitations remain in advanced wound-management reasoning, procedural planning, and autonomous clinical reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。