用图结构编码多模态脑影像,提升生成与重建效果
Latent graph encoding of multimodal neuroimaging features with generative AI architectures

- 将功能连接图转化为低维潜在空间的图结构编码
- 所提gMMVAE模型在生成质量、重建精度上均更优
- 适合脑科学中多模态影像分析的研究者使用
尽管生成模型可对复杂神经影像数据进行特征生成与重建,但设计合适的架构框架及编码与潜在空间处理机制对于研究大脑结构与功能特性至关重要。本文通过系统评估编码策略、潜在空间多模态融合方式以及生成模型选择,构建了针对结构与功能磁共振成像(MRI)特征的多模态生成框架。基于大规模神经影像数据集中的灰质体积(GMV)与静态功能网络连接(sFNC)特征,我们对比了变分自编码器(VAEs)、Transformer、生成对抗网络(GANs)和扩散模型等架构。采用模态感知图编码将功能连接嵌入低维潜在空间的模型,优于向量编码或直接数据空间方法。所提出的多模态图变分自编码器(gMMVAE)在生成保真度、重建质量、效率与潜在空间可区分性等多个指标上超越其他生成变体,展现出在鲁棒多模态神经影像分析中的潜力。
原文摘要 · Abstract (English)
While generative models enable encoding of complex neuroimaging data for feature generation and reconstruction, developing optimal architectural frameworks with appropriate encoding and latent space processes is crucial for studying structural and functional properties of the brain. We design a multimodal generative framework for structural and functional magnetic resonance imaging (MRI) features through systematic evaluation of encoding strategies, latent multimodal fusion, and generative model selection. Using structural gray matter volume (GMV) and static functional network connectivity (sFNC) features from a large neuroimaging dataset, we analyze generative frameworks involving variational autoencoders (VAEs), transformers, generative adversarial networks (GANs), and diffusion models. Architectures that employ modality-aware graph encoding of functional connectivity into a lower-dimensional latent space outperform vectorized encoders or direct data space approaches. The proposed multimodal graph VAE (gMMVAE) surpasses alternative generative variants across multiple metrics for generation fidelity, reconstruction quality, efficiency, and latent space discriminability, highlighting its potential for robust multimodal neuroimaging analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。