首个联合生成MRI与临床数据的扩散模型,助力医疗数字孪生。
Multimodal synthesis of MRI and tabular data with diffusion in a joint latent space via cross-attention

- 通过跨注意力机制在共享潜空间中融合多模态数据
- 生成的MRI结构合理,且与合成表型数据一致
- 适合医疗数据合成与数字孪生研究者参考
我们提出一种多模态潜在扩散模型,通过跨注意力机制在共享潜空间中联合生成体积分辨率磁共振成像(MRI)和临床表型数据。该方法实现对MRI与表型模态的协同表示学习,用于生成建模。模型采用变分自编码器融合双模态数据,再进行基于扩散的合成,使用独立解码器分别重建MRI与表型数据。我们在德国国家队列(NAKO Gesundheitsstudie)数据集上评估,该数据集包含超过10,000名参与者,其临床特征包括年龄、性别、身体测量值及种族等。生成的MRI体积具有解剖合理性,且身体成分与合成表型属性一致。定量评估显示,使用弗雷歇距离与精确-召回指标均验证了高质量图像生成。在表型模态上,本模型优于CTGAN,在标准评估指标上表现接近TVAE,达到与现有单模态基线相当的性能。据作者所知,这是首个在单一潜空间扩散框架中联合建模MRI与混合类型表型数据的工作,为生成一致的合成多模态患者数据提供了可行性证明,并契合医疗数字孪生的发展目标。
原文摘要 · Abstract (English)
We propose a multimodal latent diffusion model that jointly synthesizes volumetric magnetic resonance imaging (MRI) and tabular clinical data within a shared latent space via cross-attention. This approach enables coherent joint representation learning of MRI and tabular modalities for generative modeling. Our model utilizes a variational autoencoder to fuse the two modalities before diffusion-based synthesis, allowing modality-appropriate reconstruction with separate decoders for MRI and tabular data. We evaluated the framework on data from the German National Cohort (NAKO Gesundheitsstudie), comprising over 10,000 participants with MRI scans and clinical tabular features such as age, sex, body measurements, and ethnicity. The generated MRI volumes exhibited anatomical plausibility and body composition consistent with the synthesized tabular attributes. Quantitative evaluation using Fréchet distance and precision-recall metrics confirmed high-fidelity image generation. In the tabular modality, our model outperformed CTGAN across standard evaluation metrics and achieved results comparable to TVAE, demonstrating competitive performance relative to established unimodal baselines. This work is, to our knowledge, the first to demonstrate the feasibility of jointly modeling MRI and mixed-type tabular data in a single latent diffusion framework, offering a proof-of-concept for generating coherent synthetic multimodal patient data and aligning with the broader goal of developing digital twins in healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。