arXiv:2608.18937cs.CLcs.AI2026-08

构建首个大规模医学多模态统一理解生成框架

MedUAG: Unified Understanding and Generation for Medical Multimodal Models

论文配图:MedUAG: Unified Understanding and Generation for Medical Multimodal Models
图 1 · 摘自论文原文
  • 构建超600万条数据的医学多模态统一数据集
  • 在12项任务中实现强于现有模型的生成性能
  • 适合医疗AI研究者和多模态系统开发者使用

近期多模态大模型正向统一理解与生成(UAG)范式演进,但其在医疗领域的应用受限于缺乏全面的训练与评估基准,以及未形成广泛验证的统一医学模型。为此,我们提出医学UAG的综合性基础:首先构建了迄今最大的医学多模态理解与生成数据集MedUAGCorpus,包含超过600万条实例,覆盖14种影像模态;其次设计了MedUAGBench,一套标准化协议下的12类多样化医学生成评估任务;最后基于这些资源,训练出端到端的统一医学模型MedUAG。大量实验证明,MedUAG在多种理解与生成任务中表现优异,建立了具有竞争力的基线,为下一代医学多模态系统铺平道路。

原文摘要 · Abstract (English)

Recent Multimodal Large Language Models (MLLMs) are rapidly evolving into unified understanding and generation (UAG) frameworks. However, extending these unified paradigms to the medical domain is hindered by: the absence of comprehensive training and evaluation benchmarks, and the lack of broadly validated unified medical model. To address these gaps, we present a comprehensive foundation for medical UAG. First, we construct MedUAGCorpus, the largest unified medical understanding and generation dataset to date, comprising over 6 million instances across 14 imaging modalities. Second, we introduce MedUAGBench, a systematic benchmark that expands medical generation evaluation to 12 diverse tasks under standardized protocols. Finally, leveraging these resources, we develop MedUAG, an end-to-end trained unified medical model. Extensive experiments demonstrate that MedUAG achieves strong performance across a wide array of understanding and generation tasks, establishing a competitive baseline and paving the way for next-generation medical multimodal systems.

医学多模态统一生成大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。