arXiv:2507.05201cs.AIcs.CL2025-07被引 437

MedGemma是医疗视觉语言基础模型,提升医学多模态理解与推理能力。

MedGemma Technical Report

  • 基于Gemma 3构建医疗多模态模型,支持图像与文本联合理解。
  • 在胸部X光分类等任务上性能提升15.5%-18.1%,接近专用模型水平。
  • 适合医疗科研与下游应用开发,尤其需少样本微调的场景。

人工智能在医疗领域潜力巨大,但受限于数据多样性、任务复杂性及隐私保护需求。具备强大医学理解与推理能力的基础模型可减少对特定任务数据的依赖,加速医疗AI发展。我们提出MedGemma,基于Gemma 3 4B和27B的医疗视觉-语言基础模型集合。MedGemma在图像与文本联合理解任务中表现优异,显著优于同规模生成模型,并接近专用模型性能,同时保持Gemma 3的通用能力。在分布外任务中,其在医学多模态问答上提升2.6%-10%,胸部X光病灶分类提升15.5%-18.1%,代理评估提升10.8%。微调后,在电子病历信息检索中错误率降低50%,在气胸分类与组织病理切片分类上达到现有先进方法水平。我们还推出MedSigLIP——基于SigLIP的医学优化视觉编码器,作为MedGemma视觉理解核心,性能可比或优于专用医学图像编码器。MedGemma系列模型(含教程与权重)可在https://goo.gle/medgemma获取。

原文摘要 · Abstract (English)

Artificial intelligence (AI) has significant potential in healthcare applications, but its training and deployment faces challenges due to healthcare's diverse data, complex tasks, and the need to preserve privacy. Foundation models that perform well on medical tasks and require less task-specific tuning data are critical to accelerate the development of healthcare AI applications. We introduce MedGemma, a collection of medical vision-language foundation models based on Gemma 3 4B and 27B. MedGemma demonstrates advanced medical understanding and reasoning on images and text, significantly exceeding the performance of similar-sized generative models and approaching the performance of task-specific models, while maintaining the general capabilities of the Gemma 3 base models. For out-of-distribution tasks, MedGemma achieves 2.6-10% improvement on medical multimodal question answering, 15.5-18.1% improvement on chest X-ray finding classification, and 10.8% improvement on agentic evaluations compared to the base models. Fine-tuning MedGemma further improves performance in subdomains, reducing errors in electronic health record information retrieval by 50% and reaching comparable performance to existing specialized state-of-the-art methods for pneumothorax classification and histopathology patch classification. We additionally introduce MedSigLIP, a medically-tuned vision encoder derived from SigLIP. MedSigLIP powers the visual understanding capabilities of MedGemma and as an encoder achieves comparable or better performance than specialized medical image encoders. Taken together, the MedGemma collection provides a strong foundation of medical image and text capabilities, with potential to significantly accelerate medical research and development of downstream applications. The MedGemma collection, including tutorials and model weights, can be found at https://goo.gle/medgemma.

医疗AI视觉语言基础模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。