用扩散模型生成受任务刺激的脑部动态影像,效果逼近真实数据。
Scalable Diffusion Transformer for Conditional 4D fMRI Synthesis
- 结合3D VQ-GAN与CNN-Transformer,实现体素级4D fMRI条件生成。
- 在HCP数据上达到0.83的任务激活图相关性与0.98的代表结构一致性。
- 适合脑科学、神经影像建模及虚拟实验研究者使用。
由于跨被试/采集的高维异质性血氧水平依赖(BOLD)动态以及缺乏神经科学验证,生成全脑4D fMRI序列仍具挑战。本文提出首个用于体素级4D fMRI条件生成的扩散变换器,融合3D VQ-GAN隐空间压缩与CNN-Transformer主干网络,并通过AdaLN-Zero和交叉注意力实现强任务条件控制。在HCP任务fMRI数据上,模型重现了任务诱发激活图,保持真实数据中观察到的跨任务表征结构一致性(RSA),实现完美条件特异性,并使感兴趣区(ROI)时间序列与典型血流动力学响应对齐。性能随规模提升而稳定改善,任务激活图相关性达0.83,RSA为0.98,全面优于U-Net基线。该工作通过将隐空间扩散与可扩展主干及强条件控制结合,为条件化4D fMRI合成提供了实用路径,有望推动虚拟实验、跨站点标准化及下游神经影像模型的数据增强应用。
原文摘要 · Abstract (English)
Generating whole-brain 4D fMRI sequences conditioned on cognitive tasks remains challenging due to the high-dimensional, heterogeneous BOLD dynamics across subjects/acquisitions and the lack of neuroscience-grounded validation. We introduce the first diffusion transformer for voxelwise 4D fMRI conditional generation, combining 3D VQ-GAN latent compression with a CNN-Transformer backbone and strong task conditioning via AdaLN-Zero and cross-attention. On HCP task fMRI, our model reproduces task-evoked activation maps, preserves the inter-task representational structure observed in real data (RSA), achieves perfect condition specificity, and aligns ROI time-courses with canonical hemodynamic responses. Performance improves predictably with scale, reaching task-evoked map correlation of 0.83 and RSA of 0.98, consistently surpassing a U-Net baseline on all metrics. By coupling latent diffusion with a scalable backbone and strong conditioning, this work establishes a practical path to conditional 4D fMRI synthesis, paving the way for future applications such as virtual experiments, cross-site harmonization, and principled augmentation for downstream neuroimaging models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。