arXiv:2512.23304cs.CVcs.AI2025-12中稿 · publication in the…

开源医学模型MedGemma在六类疾病图像诊断中表现优于GPT-4。

MedGemma vs GPT-4: Open-Source and Proprietary Zero-shot Medical Disease Classification from Images

  • 用LoRA微调开源医学模型MedGemma-4b-it
  • 平均准确率达80.37%,显著高于GPT-4的69.58%
  • 在癌症和肺炎等高风险任务中敏感性更强

多模态大语言模型为医学影像分析引入新范式,通过临床知识解读影像实现疾病分类。本研究对比了专用开源模型MedGemma与专有大模型GPT-4在六类疾病诊断中的表现。经低秩适配(LoRA)微调的MedGemma-4b-it模型在测试集上达到80.37%的平均准确率,优于未微调的GPT-4的69.58%。此外,MedGemma在癌症与肺炎等高风险临床任务中展现出更高敏感性。混淆矩阵与分类报告的定量分析揭示了各模型在所有类别中的性能差异。结果表明,领域特定微调对降低临床应用中的幻觉至关重要,使MedGemma成为复杂、基于证据的医学推理的理想工具。

原文摘要 · Abstract (English)

Multimodal Large Language Models (LLMs) introduce an emerging paradigm for medical imaging by interpreting scans through the lens of extensive clinical knowledge, offering a transformative approach to disease classification. This study presents a critical comparison between two fundamentally different AI architectures: the specialized open-source agent MedGemma and the proprietary large multimodal model GPT-4 for diagnosing six different diseases. The MedGemma-4b-it model, fine-tuned using Low-Rank Adaptation (LoRA), demonstrated superior diagnostic capability by achieving a mean test accuracy of 80.37% compared to 69.58% for the untuned GPT-4. Furthermore, MedGemma exhibited notably higher sensitivity in high-stakes clinical tasks, such as cancer and pneumonia detection. Quantitative analysis via confusion matrices and classification reports provides comprehensive insights into model performance across all categories. These results emphasize that domain-specific fine-tuning is essential for minimizing hallucinations in clinical implementation, positioning MedGemma as a sophisticated tool for complex, evidence-based medical reasoning.

医学图像多模态模型零样本分类开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。