arXiv:2411.07121cs.CV2024-11被引 11

用全脑fMRI模型解码视觉体验,准确率提升43%。

Decoding Visual Experience and Mapping Semantics through Whole-Brain Analysis Using fMRI Foundation Models

  • 结合Transformer和生成模型,通过图像-fMRI对比学习训练全脑解码器。
  • 在BOLD5000数据集上,语义解码准确率比现有方法高43%。
  • 发现默认模式网络对视觉意义理解至关重要,适合认知神经科学研究者。

神经解码旨在揭示脑活动与不同刺激之间的对应关系,是认知科学的核心目标。过去三十年,功能磁共振成像(fMRI)与机器学习的发展显著提升了将视觉刺激映射到视觉皮层的能力。近年来,研究已拓展至利用新技术解码语言、记忆等更复杂过程,以应对更大变异性和提升信号精度。我们认为,'看见'不仅限于视觉皮层的映射,而是涉及全脑活动,因不同场景可引发情绪与认知状态变化。本文开发算法,通过个体观看视觉刺激时的全脑激活图增强对视觉过程的理解。我们采用基于Transformer的大规模fMRI编码器及在公开数据集上预训练的图像生成模型(编码器与解码器),并通过图像-fMRI对比学习进行微调。所提模型可在整个大脑皮层实现视觉体验解码,突破传统仅依赖视觉皮层的局限。在公开数据集BOLD5000上,与现有最优方法相比,语义解码准确率提升43%。网络消融分析表明,除视觉皮层外,默认模式网络对刺激解码贡献显著,符合其在意义建构与语义处理中的作用。

原文摘要 · Abstract (English)

Neural decoding, the process of understanding how brain activity corresponds to different stimuli, has been a primary objective in cognitive sciences. Over the past three decades, advances in functional Magnetic Resonance Imaging (fMRI) and machine learning have greatly improved our ability to map visual stimuli to brain activity, especially in the visual cortex. Concurrently, research has expanded to decode more complex processes, such as language and memory across the whole brain, using techniques to handle greater variability and improve signal accuracy. We argue that "seeing" involves more than just mapping visual stimuli onto the visual cortex; it engages the entire brain, as various emotions and cognitive states can emerge from observing different scenes. In this paper, we develop algorithms to enhance our understanding of visual processes by incorporating whole-brain activation maps while individuals are exposed to visual stimuli. We utilize transformer-based large-scale fMRI encoders and Image generative models (encoders & decoders) pre-trained on large public datasets, which are then fine-tuned through Image-fMRI contrastive learning. Our models can decode visual experience across the entire cerebral cortex, surpassing the traditional confines of the visual cortex. Using a public dataset (BOLD5000), we first compare our method with state-of-the-art approaches for decoding visual processing and show improved predictive semantic accuracy by 43%. A network ablation analysis suggests that, beyond the visual cortex, the default mode network contributes significantly to stimulus decoding, in line with the proposed role of this network in sense-making and semantic processing.

脑机接口fMRI解码生成模型全脑分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。