arXiv:2501.16471cs.LGcs.AI2025-01ICLR被引 6

用表面视觉变压器实现跨个体电影观看脑活动解码

SIM: Surface-based fMRI Analysis for Inter-Subject Multimodal Decoding from Movie-Watching Experiments

  • 基于脑皮层表面构建动态视觉变换模型,捕捉功能网络拓扑
  • 在174人7T fMRI数据上实现未训练人物/影片的准确解码
  • 适合开发个性化脑机接口与神经反馈系统

当前脑解码与编码框架通常在同一批数据上训练和测试,限制了其在脑机接口(BCI)或神经反馈中的应用。若能跨个体整合经验以模拟训练中未见的刺激,则更具实用性。主要障碍在于个体间皮层组织的差异性,导致皮层信号难以对齐或比较。本文提出表面视觉变换器,将皮层网络拓扑及其交互关系建模为表面移动图像,构建可泛化的皮层功能动态模型。该模型结合音频、视频与fMRI三模态自监督对比(CLIP)对齐,实现从脑活动模式中检索视觉/听觉刺激(反之亦然)。在包含174名健康被试的HCP电影观看任务7T任务fMRI数据上验证:即使面对训练中未见的个体和影片,仍可仅凭脑活动准确识别正在观看的片段。注意力图分析显示,模型捕捉到反映语义与视觉系统的个体化脑活动模式,为未来个性化脑功能模拟开辟可能。代码与预训练模型将开源,训练数据可申请获取。

原文摘要 · Abstract (English)

Current AI frameworks for brain decoding and encoding, typically train and test models within the same datasets. This limits their utility for brain computer interfaces (BCI) or neurofeedback, for which it would be useful to pool experiences across individuals to better simulate stimuli not sampled during training. A key obstacle to model generalisation is the degree of variability of inter-subject cortical organisation, which makes it difficult to align or compare cortical signals across participants. In this paper we address this through the use of surface vision transformers, which build a generalisable model of cortical functional dynamics, through encoding the topography of cortical networks and their interactions as a moving image across a surface. This is then combined with tri-modal self-supervised contrastive (CLIP) alignment of audio, video, and fMRI modalities to enable the retrieval of visual and auditory stimuli from patterns of cortical activity (and vice-versa). We validate our approach on 7T task-fMRI data from 174 healthy participants engaged in the movie-watching experiment from the Human Connectome Project (HCP). Results show that it is possible to detect which movie clips an individual is watching purely from their brain activity, even for individuals and movies not seen during training. Further analysis of attention maps reveals that our model captures individual patterns of brain activity that reflect semantic and visual systems. This opens the door to future personalised simulations of brain function. Code & pre-trained models will be made available at https://github.com/metrics-lab/sim, processed data for training will be available upon request at https://gin.g-node.org/Sdahan30/sim.

脑解码多模态表面建模跨个体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。