arXiv:2602.23410cs.LGcs.AI2026-02

统一处理脑电、脑磁和功能核磁,提升神经影像分析效果

Brain-OF: An Omnifunctional Foundation Model for fMRI, EEG and MEG

  • 用统一框架同时处理fMRI、EEG、MEG三种脑成像数据
  • 在40个数据集上预训练,多任务表现优于单一模态模型
  • 适合需要融合多种脑信号的研究者,如脑机接口与认知研究

脑基础模型已在众多神经科学任务中取得显著进展,但多数现有模型仅限于单一功能模态,难以利用不同神经成像技术间的互补时空动态和集体数据规模。这一局限主要源于各模态间严重的语义异质性与分辨率差异。为此,我们提出Brain-OF,一个在fMRI、EEG和MEG上联合预训练的全功能脑基础模型,可在统一框架内处理单模态与多模态输入。为解决异构时空分辨率问题,引入任意分辨率神经信号采样器,将多样脑信号映射至共享语义空间。为应对语义偏移,Brain-OF骨干网络结合DINT注意力与稀疏专家混合机制,其中共享专家捕捉模态不变表示,路由专家专注模态特异性语义。此外,通过自监督学习显式内化神经活动特征,提出掩码时频建模,一种在时域与频域联合重建脑信号的双域预训练目标。Brain-OF在包含约40个数据集的大规模语料上预训练,展现出在多样化下游任务中的优越性能,凸显联合多模态整合与双域预训练的优势。

原文摘要 · Abstract (English)

Brain foundation models have achieved remarkable advances across a wide range of neuroscience tasks. However, most existing models are limited to a single functional modality, restricting their ability to exploit complementary spatiotemporal dynamics and the collective data scale across different neuroimaging techniques. This limitation largely arises from severe semantic heterogeneity and resolution discrepancies among modalities. To address these challenges, we propose Brain-OF, an omnifunctional brain foundation model jointly pretrained on fMRI, EEG and MEG, capable of handling both unimodal and multimodal inputs within a unified framework. To reconcile heterogeneous spatiotemporal resolutions, we introduce the Any-Resolution Neural Signal Sampler, which projects diverse brain signals into a shared semantic space. To further manage semantic shifts, the Brain-OF backbone integrates DINT attention with a Sparse Mixture of Experts, where shared experts capture modality-invariant representations and routed experts specialize in modality-specific semantics. Furthermore, to explicitly internalize the characteristics of neural activity through self-supervised learning, we propose Masked Temporal-Frequency Modeling, a dual-domain pretraining objective that jointly reconstructs brain signals in both the time and frequency domains. Brain-OF is pretrained on a large-scale corpus comprising around 40 datasets and demonstrates superior performance across diverse downstream tasks, highlighting the benefits of joint multimodal integration and dual-domain pretraining.

脑机接口多模态基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。