arXiv:2508.08270cs.LGcs.AI2025-08

Doctor Sun是专为医学设计的双语多模态大模型,提升医图像与文本理解能力。

Doctor Sun: A Bilingual Multimodal Large Language Model for Biomedical AI

  • 融合视觉编码器与医学大模型,分两阶段训练对齐图文特征。
  • 在多个医学数据集上实现跨模态理解,支持中英文医疗任务。
  • 开源了SunMed-VL双语医学多模态数据集及完整资源,推动研究开放。

大型多模态模型(LMMs)在病理分析、影像报告生成和生物医学辅助等领域展现出巨大潜力。然而,现有医学多模态AI多基于通用基础大语言模型,受限于医疗训练数据,难以深入理解复杂医学概念;近期受LLaVA启发的医学多模态模型也难以有效捕捉文本与图像间的细微关联。为此,我们提出Doctor Sun,一个专为医学设计的大型生成式多模态模型,可编码、融合并解释文本与图像等多种生物医学数据。具体而言,Doctor Sun结合预训练视觉编码器与医学专用大模型,在多个医学数据集上进行两阶段训练,聚焦特征对齐与指令微调。此外,我们发布了SunMed-VL——一个覆盖广泛的双语医学多模态数据集,配套提供全部模型、代码与资源,免费支持生物医学多模态研究发展。

原文摘要 · Abstract (English)

Large multimodal models (LMMs) have demonstrated significant potential in providing innovative solutions for various biomedical tasks, including pathology analysis, radiology report generation, and biomedical assistance. However, the existing multimodal biomedical AI is typically based on foundation LLMs, thus hindering the understanding of intricate medical concepts with limited medical training data. Moreover, recent LLaVA-induced medical LMMs struggle to effectively capture the intricate relationship between the texts and the images. Therefore, we introduce Doctor Sun, a large multimodal generative model specialized in medicine, developed to encode, integrate, and interpret diverse biomedical data modalities such as text and images. In particular, Doctor Sun integrates a pre-trained vision encoder with a medical LLM and conducts two-stage training on various medical datasets, focusing on feature alignment and instruction tuning. Moreover, we release SunMed-VL, a wide-range bilingual medical multimodal dataset, along with all associated models, code, and resources, to freely support the advancement of biomedical multimodal research.

多模态模型医学AI双语大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。