实测发现,主流多模态模型在真实皮肤科诊疗中表现远逊于公开数据集。
Are Multimodal LLMs Ready for Clinical Dermatology? A Real-World Evaluation in Dermatology
- 在5811例真实临床病例上评估5个模型,对比公开数据集表现
- 仅用图像时诊断准确率最高降至13.35%,结合病史后提升至28.75%
- 模型对病史信息敏感,错误输入易导致误诊,尚不适合临床部署
多模态大语言模型(MLLMs)在公开皮肤病学基准测试中表现良好,但其性能是否适用于真实临床场景仍不确定。本研究在三个公开皮肤病数据集和一个包含5,811例病例、46,405张临床图像的多中心医院回顾性会诊队列上,评估了四个开源模型(InternVL-Chat v1.5、LLaVA-Med v1.5、SkinGPT4、MedGemma-4B-Instruct)和一个商用模型(GPT-4.1)。任务包括鉴别诊断生成与基于严重程度的分诊。在公开数据集上,最优开源模型的前3名诊断准确率为26.55%,GPT-4.1达42.25%;而在真实会诊案例中仅用图像时,开源模型准确率降至1.50%-13.35%,GPT-4.1为24.65%。加入临床背景信息后,开源模型最高提升至28.75%,GPT-4.1达38.93%,但模型输出对不完整或错误的病史极为敏感。在分诊任务中,模型敏感度超过60%,显示筛查潜力,但可靠性不足,难以直接用于临床决策。结果表明,当前皮肤病学多模态模型的真实临床能力被显著高估。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have demonstrated promise on publicly available dermatology benchmarks. However, benchmark performance may not generalize to real-world dermatologic decision-making. To quantify this benchmark-to-bedside gap, we evaluated four open-weight MLLMs (InternVL-Chat v1.5, LLaVA-Med v1.5, SkinGPT4 and MedGemma-4B-Instruct) and one commercial MLLM (GPT-4.1) across three publicly available dermatology datasets and a retrospective multi-site hospital-based dermatology consultation cohort comprising 5,811 cases and 46,405 clinical images. Models were evaluated on two clinically relevant tasks: differential diagnosis generation and severity-based triage. Diagnostic performance was modest on public datasets and declined substantially in the real-world cohort. On public benchmarks, top-3 diagnostic accuracy reached 26.55% for the best open-weight model and 42.25% for GPT-4.1. On real-world consultation cases using images alone, top-3 diagnostic accuracy fell to 1.50%-13.35% among open-weight models and 24.65% for GPT-4.1. Incorporating clinical context improved performance across all models, increasing top-3 diagnostic accuracy up to 28.75% among open-weight models and 38.93% for GPT-4.1. However, model outputs were highly sensitive to incomplete or erroneous consultation context. For severity-based triage, models achieved moderate sensitivity (above 60%), suggesting potential utility for screening but insufficient reliability for clinical deployment. These findings demonstrate that benchmark performance substantially overestimates the real-world clinical capability of current dermatology MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。