MedViLaM用同一模型处理医学图文,实现跨任务通用与可解释。
MedViLaM: A multimodal large language model with advanced generalizability and explainability for medical data understanding and generation
- 统一模型同时理解医学文本与影像,共享参数提升泛化能力。
- 在多任务基准上表现超越多数现有模型,零样本迁移效果显著。
- 适合临床辅助诊断、医学报告生成等场景,医生可理解模型推理过程。
医学数据天然多模态且多任务,涵盖文本与影像等多种形式。然而,当前多数医疗模型为单模态单一任务,泛化与可解释性不足。本研究提出MedViLaM,一种面向医疗数据的统一视觉-语言通用模型,能灵活编码与解析临床文本与医学影像,仅使用一套模型权重。为支持此类多任务模型构建,我们构建了MultiMedBench——一个包含连续问答、多标签疾病分类、疾病定位、放射科报告生成与摘要等任务的综合性预训练数据集与评估基准。MedViLaM在所有MultiMedBench任务中均表现优异,常显著超越其他通用模型。此外,模型展现出对新医学概念与任务的零样本泛化能力、跨任务有效迁移学习能力,以及零样本医学推理的涌现能力。
原文摘要 · Abstract (English)
Medicine is inherently multimodal and multitask, with diverse data modalities spanning text, imaging. However, most models in medical field are unimodal single tasks and lack good generalizability and explainability. In this study, we introduce MedViLaM, a unified vision-language model towards a generalist model for medical data that can flexibly encode and interpret various forms of medical data, including clinical language and imaging, all using the same set of model weights. To facilitate the creation of such multi-task model, we have curated MultiMedBench, a comprehensive pretaining dataset and benchmark consisting of several distinct tasks, i.e., continuous question-answering, multi-label disease classification, disease localization, generation and summarization of radiology reports. MedViLaM demonstrates strong performance across all MultiMedBench tasks, frequently outpacing other generalist models by a significant margin. Additionally, we present instances of zero-shot generalization to new medical concepts and tasks, effective transfer learning across different tasks, and the emergence of zero-shot medical reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。