让医学多模态理解与生成相互促进,提升图像合成效果。
SynerMedGen: Synergizing Medical Multimodal Understanding with Generation via Task Alignment

- 通过任务对齐机制,让理解任务为生成服务。
- 零样本下在22个任务上表现优异,跨数据集泛化能力强。
- 适合医学图像生成与多模态模型研究者使用。
统一医学多模态理解与生成是前沿方向,但现有模型常将两者割裂。本文提出生成对齐的理解原则,构建SynerMedGen框架,引入三个生成对齐的理解任务和两阶段训练策略,将理解阶段学到的生成有益表征用于医学图像合成。仅通过理解训练,该模型在22项医学图像合成任务中即实现强零样本性能,并在未见数据集上表现出稳健泛化能力。结合生成训练后,其性能持续优于最先进的专用合成模型及近期统一模型。研究还发布了包含100万组配对合成样本和200万条生成衍生理解实例的大规模数据集SynerMed,以支持理解-生成协同研究。项目地址:https://github.com/piooip/SynerMedGen。
原文摘要 · Abstract (English)
Unifying multimodal understanding and generation is a compelling frontier that is beginning to emerge in the medical field. However, the limited existing unified medical models typically treat understanding and generation as disjoint objectives, lacking a meaningful functional synergy. In this work, we identify and address a critical question in unified medical modeling: what form of understanding truly benefits generation. We present SynerMedGen, a unified framework built on the proposed principle of generation-aligned understanding, which synergizes understanding objectives with generation tasks via task alignment. SynerMedGen introduces three generation-aligned understanding tasks and a two-stage training strategy that transfers generation-beneficial representations learned during understanding training to medical image synthesis. Remarkably, even with understanding training alone, our SynerMedGen achieves strong zero-shot performance across 22 medical image synthesis tasks and demonstrates robust generalization to unseen datasets. When combined with generation training, SynerMedGen consistently outperforms state-of-the-art specialized medical image synthesis models as well as recent unified medical models. We also release a large-scale dataset named SynerMed consisting of 1M paired synthesis samples and 2M generation-derived understanding instances to support further research on understanding-generation synergy. Our project can be accessed at https://github.com/piooip/SynerMedGen.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。