arXiv:2603.00842cs.CL2026-03

200亿参数的开源医学多模态模型,支持本地部署且性能超越更大模型。

MedGPT-oss: Training a General-Purpose Vision-Language Model for Biomedicine

  • 用三阶段训练流程融合语言与视觉模块,避免复杂架构依赖。
  • 在跨域多模态推理和临床文本任务上优于更大规模开源模型。
  • 适配普通显卡,适合注重隐私保护的医疗机构研究使用。

生物医学多模态助手有望统一放射、病理与临床文本推理,但高性能系统多为闭源或计算成本过高,难以满足患者隐私与受保护健康信息(PHI)合规要求。我们提出MEDGPT-OSS,一个200亿参数的开源权重通用视觉-语言模型,旨在推动临床AI的开放研究。不同于依赖复杂架构的方法,MEDGPT-OSS通过优化的三阶段训练流程,将GPT-oss语言主干与视觉前端结合。通过严格的数据筛选与长上下文多模态对齐,逐步进行领域适应,成功缩小能力差距。该模型在出分布(OOD)多模态推理及复杂纯文本临床任务中表现优于更大的开源医学模型。通过单一指令跟随接口统一多种模态,保持参数高效,完全兼容消费级GPU。我们公开完整训练方案、开源权重检查点及严谨评估工具包,为隐私保护、机构定制的临床AI研究提供可验证基础。

原文摘要 · Abstract (English)

Biomedical multimodal assistants have the potential to unify radiology, pathology, and clinical-text reasoning, yet a critical deployment gap remains: top-performing systems are either closed-source or computationally prohibitive, precluding the on-premises deployment required for patient privacy and PHI compliance. We introduce MEDGPT-OSS, an open-weight, 20B-parameter generalist vision-language model designed to facilitate open research in clinical AI. Rather than relying on architectural complexity, MEDGPT-OSS pairs the GPT-oss language backbone with a visual front-end via a optimized, three-stage training curriculum. By progressively domain-adapting these modules through rigorous data curation and long-context multimodal alignment, we demonstrate that a 20B model can bridge the capacity gap. It successfully outperforms larger open medical models on out-of-distribution (OOD) multimodal reasoning and complex text-only clinical tasks. By unifying diverse modalities under a single instruction-following interface, MEDGPT-OSS maintains a parameter-efficient footprint fully compatible with commodity GPUs. We release the complete training recipe, open-weight checkpoints, and a rigorous evaluation harness to serve as a verifiable foundation for privacy-preserving, institution-specific clinical AI research.

多模态医学AI开源模型视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。