arXiv:2507.22229cs.LG2025-07被引 25

首个跨模态预测全脑活动的深度模型,提升认知建模统一性

TRIBE: TRImodal Brain Encoder for whole-brain fMRI response prediction

  • 融合文本、音频、视频预训练表示,用Transformer建模时序变化
  • 在Algonauts 2025竞赛中以显著优势夺冠,高阶联合皮层预测更精准
  • 适合脑科学、多模态认知建模研究者,推动统一认知模型发展

传统神经科学倾向于分领域研究,聚焦单一模态、任务或脑区,虽成果丰硕,却阻碍了认知统一模型的发展。本文提出TRIBE,首个可跨模态、跨皮层区域及跨个体预测脑响应的深度神经网络。通过整合文本、音频与视频基础模型的预训练表征,并利用Transformer处理其时变特性,该模型能精确建模对视频刺激的时空fMRI响应,在Algonauts 2025脑编码竞赛中以显著优势夺得第一名。消融实验表明,单模态模型仅能可靠预测对应皮层网络(如视觉或听觉网络),但在高阶联合皮层上均被本多模态模型超越。当前应用于感知与理解任务,为构建人类大脑表征的整合模型铺平道路。代码已开源:https://github.com/facebookresearch/algonauts-2025。

原文摘要 · Abstract (English)

Historically, neuroscience has progressed by fragmenting into specialized domains, each focusing on isolated modalities, tasks, or brain regions. While fruitful, this approach hinders the development of a unified model of cognition. Here, we introduce TRIBE, the first deep neural network trained to predict brain responses to stimuli across multiple modalities, cortical areas and individuals. By combining the pretrained representations of text, audio and video foundational models and handling their time-evolving nature with a transformer, our model can precisely model the spatial and temporal fMRI responses to videos, achieving the first place in the Algonauts 2025 brain encoding competition with a significant margin over competitors. Ablations show that while unimodal models can reliably predict their corresponding cortical networks (e.g. visual or auditory networks), they are systematically outperformed by our multimodal model in high-level associative cortices. Currently applied to perception and comprehension, our approach paves the way towards building an integrative model of representations in the human brain. Our code is available at https://github.com/facebookresearch/algonauts-2025.

脑科学多模态fMRITransformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。