用元数据控制的扩散模型,让脑影像模型学会跨状态通用表征。
Brain-DiT: A Universal Multi-state fMRI Foundation Model with Metadata-Conditioned Pretraining

- 通过元数据条件扩散训练,学习多尺度脑活动表征。
- 在7个下游任务中,生成式预训练性能优于重建与对齐方法。
- 不同任务偏好不同层次结构:分类重全局语义,预测重局部细节。
当前功能磁共振(fMRI)基础模型通常依赖有限脑状态和不匹配的预训练任务,限制了跨脑状态的通用表征学习能力。本文提出Brain-DiT,一个在24个数据集共349,898个会话上预训练的通用多状态fMRI基础模型,覆盖静息态、任务态、自然刺激态、疾病态及睡眠态。不同于以往在原始信号或潜在空间进行掩码重建的模型,Brain-DiT采用元数据条件扩散预训练,结合扩散变换器(DiT),可同时捕捉精细功能结构与全局语义。在7个下游任务上的广泛评估与消融实验表明,基于扩散的生成式预训练是比重建或对齐更优的代理任务;元数据条件化进一步提升性能,有效分离内在神经动力学与群体差异。此外,不同任务对表征尺度有不同偏好:ADNI分类更依赖全局语义,而年龄/性别预测则更依赖局部精细结构。
原文摘要 · Abstract (English)
Current fMRI foundation models primarily rely on a limited range of brain states and mismatched pretraining tasks, restricting their ability to learn generalized representations across diverse brain states. We present Brain-DiT, a universal multi-state fMRI foundation model pretrained on 349,898 sessions from 24 datasets spanning resting, task, naturalistic, disease, and sleep states. Unlike prior fMRI foundation models that rely on masked reconstruction in the raw-signal space or a latent space, Brain-DiT adopts metadata-conditioned diffusion pretraining with a Diffusion Transformer (DiT), enabling the model to learn multi-scale representations that capture both fine-grained functional structure and global semantics. Across extensive evaluations and ablations on 7 downstream tasks, we find consistent evidence that diffusion-based generative pretraining is a stronger proxy than reconstruction or alignment, with metadata-conditioned pretraining further improving downstream performance by disentangling intrinsic neural dynamics from population-level variability. We also observe that downstream tasks exhibit distinct preferences for representational scale: ADNI classification benefits more from global semantic representations, whereas age/sex prediction comparatively relies more on fine-grained local structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。