首个统一医学多模态生成的离散扩散模型,实现影像、文本、报告联合生成。
Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
- 用多模态大语言模型作扩散主干,共享跨模态概率空间。
- 生成图像与报告的FID低于25,报告METEOR达0.26以上。
- 适合医疗生成、临床辅助诊断等需要多模态协同的场景。
当前生成式医学模型受限于模态专用设计,难以整合影像、病理与临床文本的互补信息,阻碍其发展为能跨生物医学数据全面学习与推理的基础模型。本文提出MeDiM,首个无模态特异性组件的医学离散扩散模型,可学习多模态共享分布。它统一了多种生成任务:图像与文本互转,以及根据提示跨域联合生成图像-报告对。基于离散扩散框架,通过共享概率空间连接视觉与语言表示。采用多模态大语言模型(MLLM)作为扩散主干,利用其先验知识与跨模态推理能力。引入两项关键设计:(1) 移除因果注意力掩码以支持双向上下文;(2) 注入连续时间步嵌入以增强扩散感知。实验显示高保真生成性能(在MIMIC-CXR上FID 16.60,PathGen上FID 24.19),报告生成准确率高(METEOR 0.2650和0.2580)。联合生成的图像-报告对显著提升下游任务表现(BLEU-1提升6.43%,BLEU-2提升18.57%,BLEU-3提升31.58%,METEOR提升4.80%),证明MeDiM能生成语义连贯且符合临床实际的多模态输出。
原文摘要 · Abstract (English)
Recent advances in generative medical models are constrained by modality-specific scenarios that hinder the integration of complementary evidence from imaging, pathology, and clinical notes. This fragmentation limits their evolution into foundation models that can learn and reason across the full spectrum of biomedical data. We propose MeDiM, the first medical discrete diffusion model that learns shared distributions across modalities without modality-specific components. MeDiM unifies multiple generative tasks: translating between images and text, and jointly producing image-report pairs across domains in response to prompts. Built on a discrete diffusion framework, MeDiM bridges vision and language representations through a shared probabilistic space. To enable unified and flexible medical generation, we employ a multimodal large language model (MLLM) as the diffusion backbone, leveraging its prior knowledge and cross-modal reasoning. Two key designs are introduced: (1) removing the causal attention mask for bidirectional context, and (2) injecting continuous timestep embeddings for diffusion awareness. Experiments demonstrate high-fidelity medical generation (FID 16.60 on MIMIC-CXR and FID 24.19 on PathGen) and accurate report generation (METEOR 0.2650 and 0.2580). Jointly generated image-report pairs further enhance downstream performance (plus6.43 percent BLEU-1, plus18.57 percent BLEU-2, plus31.58 percent BLEU-3, plus4.80 percent METEOR), showing that MeDiM supports coherent and clinically grounded multimodal outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。