arXiv:2605.04326q-bio.NCcs.LG2026-05被引 17

用多模态AI模型预测人脑反应,实现虚拟神经科学实验

A foundation model of vision, audition, and language for in-silico neuroscience

论文配图:A foundation model of vision, audition, and language for in-silico neuroscience
图 1 · 摘自论文原文
  • 构建三模态基础模型,统一处理视频、音频和语言输入
  • 在超1000小时数据上训练,对新刺激和受试者预测准确率显著提升
  • 可模拟经典神经实验,揭示多感官整合的精细脑区分布

认知神经科学长期依赖针对特定实验范式的专用模型,难以形成统一的认知模型。本文提出TRIBE v2,一个融合视觉、听觉与语言的三模态基础模型,能够预测多种自然情境和实验条件下的真实人脑活动。基于覆盖720名受试者、超过1000小时的fMRI数据集,该模型在新刺激、新任务和新受试者上的高分辨率脑响应预测表现远超传统线性编码模型,准确率提升数倍。关键的是,TRIBE v2支持虚拟神经实验:在经典视觉与神经语言学范式中复现了数十年实证研究的关键结果。通过提取可解释的潜在特征,模型揭示了多感官整合的精细脑区拓扑结构。这些成果确立了人工智能作为探索人类大脑功能组织统一框架的潜力。

原文摘要 · Abstract (English)

Cognitive neuroscience is fragmented into specialized models, each tailored to specific experimental paradigms, hence preventing a unified model of cognition in the human brain. Here, we introduce TRIBE v2, a tri-modal (video, audio and language) foundation model capable of predicting human brain activity in a variety of naturalistic and experimental conditions. Leveraging a unified dataset of over 1,000 hours of fMRI across 720 subjects, we demonstrate that our model accurately predicts high-resolution brain responses for novel stimuli, tasks and subjects, superseding traditional linear encoding models, delivering several-fold improvements in accuracy. Critically, TRIBE v2 enables in silico experimentation: tested on seminal visual and neuro-linguistic paradigms, it recovers a variety of results established by decades of empirical research. Finally, by extracting interpretable latent features, TRIBE v2 reveals the fine-grained topography of multisensory integration. These results establish artificial intelligence as a unifying framework for exploring the functional organization of the human brain.

脑科学多模态AI建模神经编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。