用文字提示生成指定类型的脑部MRI,支持零样本和有监督生成。
Towards General Text-guided Image Synthesis for Customized Multimodal Brain MRI Generation
- 基于文本提示控制生成,统一模型处理多种脑MRI模态。
- 在13个中心的31,407例3D数据上训练,实现高精度生成。
- 适合临床科研与大规模脑病筛查,无需额外采集扫描。
多模态脑磁共振成像在神经科学与神经病学中至关重要,但受扫描仪可及性和采集时间长的限制,多模态MRI并不常见。现有合成方法通常针对特定任务在独立数据集上训练,导致在新数据集和任务上表现不佳。本文提出TUMSyn,一个文本引导的通用脑MRI图像合成模型,可基于常规扫描和文本提示灵活生成带有指定成像元数据的脑部MRI图像。为确保生成精度、泛化性与多样性,我们构建了包含31,407例3D图像、覆盖7种MRI模态的脑部MRI数据库,来自13个中心。通过对比学习预训练专用文本编码器,有效实现文本对图像合成的控制。在多个数据集上的实验及医生评估表明,TUMSyn可在监督与零样本场景下生成具有临床意义的脑部MRI图像。因此,TUMSyn可结合已有扫描,助力大规模基于MRI的脑疾病筛查与诊断。
原文摘要 · Abstract (English)
Multimodal brain magnetic resonance (MR) imaging is indispensable in neuroscience and neurology. However, due to the accessibility of MRI scanners and their lengthy acquisition time, multimodal MR images are not commonly available. Current MR image synthesis approaches are typically trained on independent datasets for specific tasks, leading to suboptimal performance when applied to novel datasets and tasks. Here, we present TUMSyn, a Text-guided Universal MR image Synthesis generalist model, which can flexibly generate brain MR images with demanded imaging metadata from routinely acquired scans guided by text prompts. To ensure TUMSyn's image synthesis precision, versatility, and generalizability, we first construct a brain MR database comprising 31,407 3D images with 7 MRI modalities from 13 centers. We then pre-train an MRI-specific text encoder using contrastive learning to effectively control MR image synthesis based on text prompts. Extensive experiments on diverse datasets and physician assessments indicate that TUMSyn can generate clinically meaningful MR images with specified imaging metadata in supervised and zero-shot scenarios. Therefore, TUMSyn can be utilized along with acquired MR scan(s) to facilitate large-scale MRI-based screening and diagnosis of brain diseases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。