arXiv:2501.04614cs.AIcs.LG2025-01被引 9

XGeM可任意生成多模态医学数据,解决真实医疗数据难获取问题。

XGeM: A Multi-Prompt Foundation Model for Multimodal Medical Data Generation

  • 用对比学习构建共享隐空间,支持任意输入模态组合生成
  • 在MIMIC-CXR上生成的胸片与报告通过专家视觉图灵测试
  • 适用于数据匿名化、类别不平衡等临床数据难题

人工智能在医学影像中前景广阔,但受限于数据稀缺、隐私问题及多模态整合困难。现有生成模型多为单模态、单向合成,难以联合生成多种模态且保持临床一致性。为此,我们提出XGeM,一个67.7亿参数的多模态生成模型,支持任意模态间的双向合成。XGeM通过对比学习构建共享隐空间,并引入新型多提示训练策略,可基于任意输入模态子集进行条件生成,适应异构临床输入,联合生成多个输出,保持语义与结构一致。我们在MIMIC-CXR数据集上与五种基线模型对比,验证其在多视角胸部X光与放射科报告生成上的性能;并通过专家放射科医生的视觉图灵测试评估生成数据的真实性和临床相关性。此外,我们展示了XGeM在数据匿名化、类别不平衡和数据稀缺等关键医疗数据挑战中的应用价值,证实其作为医学数据合成基础模型的潜力。项目页面:https://cosbidev.github.io/XGeM/

原文摘要 · Abstract (English)

The adoption of Artificial Intelligence in medical imaging holds great promise, yet it remains hindered by challenges such as data scarcity, privacy concerns, and the need for robust multimodal integration. While recent advances in generative modeling have enabled high-quality synthetic data generation, existing approaches are often limited to unimodal, unidirectional synthesis and therefore lack the ability to jointly synthesize multiple modalities while preserving clinical consistency. To address this challenge, we introduce XGeM, a 6.77-billion-parameter multimodal generative model designed to support flexible, any-to-any synthesis between medical data modalities. XGeM constructs a shared latent space via contrastive learning and introduces a novel Multi-Prompt Training strategy, enabling conditioning on arbitrary subsets of input modalities. This design allows the model to adapt to heterogeneous clinical inputs and generate multiple outputs jointly, preserving both semantic and structural coherence. We extensively validate XGeM: first we benchmark it against five competitors on the MIMIC-CXR dataset, a state-of-the-art dataset for multi-view Chest X-ray and radiological report generation. Secondly, we perform a Visual Turing Test with expert radiologists to assess the realism and clinical relevance of the generated data, ensuring alignment with real-world scenarios. Finally, we show how XGeM can support key medical data challenges such as anonymization, class imbalance, and data scarcity, underscoring its utility as a foundation model for medical data synthesis. Project page is at https://cosbidev.github.io/XGeM/.

多模态生成医学影像生成模型数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。