用平面对映提升视觉变压器在功能MRI的性能。
Scaling Vision Transformers for Functional MRI with Flat Maps
- 将3D脑部影像转为2D平面图,适配视觉变压器架构
- 在认知状态解码任务中,模型表现远超已有方法
- 发布首个开放评估套件,支持模型公平对比
我们研究自监督基础模型在功能磁共振成像(fMRI)中的训练问题。主要贡献包括:(1) 提出新模型家族CortexMAE,基于2100小时公开fMRI数据,采用掩码自编码器框架训练;(2) 发布首个开放评估套件Brainmarks,用于fMRI基础模型评测。核心创新在于:将3D fMRI体积通过皮层平展投影转换为2D地图,直接对比平展图、脑区划分和体积表示。结果表明,平展图总体表现最优。我们首次开展系统性缩放分析,发现严格幂律缩放关系,但存在上限。使用Brainmarks进行受控基准测试:在个体特征预测任务中,未见任一模型显著超越现有水平,且所有模型均难胜过简单功能连接基线;而在认知状态解码任务中,表现更稳健,我们的CortexMAE家族大幅优于先前模型。代码、模型与数据集已开源至https://github.com/MedARC-AI/CortexMAE和https://github.com/MedARC-AI/Brainmarks。
原文摘要 · Abstract (English)
We study the problem of training self-supervised foundation models for functional MRI. Our main contributions are: (1) we introduce a new model family (CortexMAE) trained using the masked autoencoder framework on 2.1K hours of open fMRI data, and (2) we release the first open evaluation suite (Brainmarks) for fMRI foundation models. Our core innovation is simple: we adapt the Vision Transformer to fMRI by first converting each 3D fMRI volume to a 2D map using a cortical flat map projection. We directly compare flat maps to both parcellation and volume-based representations. While each has its advantages, flat maps generally perform best. We perform the first systematic scaling analysis for fMRI and observe strict power law scaling, albeit with limits. Finally, we use Brainmarks to do controlled benchmark comparisons. On subject-level trait prediction, we report a challenging null result: no single model achieves clear state-of-the-art performance. Moreover, all models struggle to outperform a simple functional connectivity baseline. On cognitive state decoding, we observe more robust performance, and in this setting our CortexMAE family outperforms prior models by a large margin. Code, models, and datasets are available at https://github.com/MedARC-AI/CortexMAE and https://github.com/MedARC-AI/Brainmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。