arXiv:2606.00689cs.CV2026-06被引 1

用小波融合技术提升多模态脑MRI合成质量,支持复杂数据场景。

Wavelet-Fusion Diffusion Model for Multimodal Brain MRI Synthesis with Modality and Metadata Conditioning

论文配图:Wavelet-Fusion Diffusion Model for Multimodal Brain MRI Synthesis with Modality and Metadata Conditioning
图 1 · 摘自论文原文
  • 基于小波融合的变分自编码器压缩体积,降低3D扩散模型计算负担。
  • 在多源异构数据上合成效果最优,分布对齐程度领先于现有方法。
  • 适合处理多模态缺失、设备差异大的真实神经影像数据集。

多模态磁共振成像可提供互补信息,支持神经影像分析与下游AI应用开发。然而,公开和整合的数据集中各模态覆盖不均,且受扫描设备、采集协议及人口临床信息差异影响,常存在数据稀疏或缺失问题。合成MRI生成可缓解这一不平衡,用于数据增强和可控合成队列构建。现有方法多在有限模态或同质群体上训练,难以适应大规模异构数据资源。扩散模型因样本保真度高、多样性好成为优选,但直接在3D体素空间采样效率低。潜空间扩散模型通过在学习的3D潜空间中生成改善实用性,但质量依赖自编码器重建精度与潜空间分布。本文提出结合小波融合变分自编码器(WF-VAE)与条件3D U-Net扩散模型,在学习潜空间中实现显式模态与元数据条件控制。所提小波融合扩散模型(WFDM)在评估中达到最佳分布对齐效果。

原文摘要 · Abstract (English)

Multimodal MRI provides complementary information for neuroimaging analysis, where different imaging modalities capture distinct anatomical, tissue, and pathological features that support the development and evaluation of downstream AI applications. Although large-scale structural MRI resources are increasingly available, their modality coverage is often uneven across public and pooled neuroimaging datasets. This uneven modality coverage is further complicated by heterogeneity across sites, scanners, and acquisition protocols, as well as demographic and clinical variables that are often sparse, inconsistently recorded, or unavailable across studies. Synthetic MRI generation can help address this imbalance by synthesizing target-modality volumes for dataset augmentation and controlled synthetic cohort creation. However, many existing MRI synthesis approaches are trained on narrow modality sets or relatively homogeneous cohorts, limiting their applicability to large pooled neuroimaging resources where modality availability, acquisition protocols, and metadata coverage vary substantially across datasets. Diffusion models have become an attractive approach for MRI synthesis because of their strong sample fidelity and diversity, but sampling directly in 3D voxel space is computationally expensive and slow at inference. Latent diffusion improves practicality by synthesizing MRI in a learned, 3D latent space, although generation quality depends on the autoencoder's reconstruction fidelity and the resulting latent distribution. Our approach combines a Wavelet-Fusion variational autoencoder (WF-VAE) latent compressor with a conditional 3D U-Net diffusion model trained in the learned latent space using explicit modality and metadata conditioning. Our proposed Wavelet-Fusion Diffusion Model (WFDM) achieved the strongest distributional alignment among the evaluated synthetic MRI generators.

脑MRI多模态生成扩散模型数据融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。