用自回归模型预测大脑在自然刺激下的动态演化,提升脑活动预测精度。
BrainVista: Modeling Naturalistic Brain Dynamics as Multimodal Next-Token Prediction
- 将多模态输入分解为独立脑网络动态,通过空间混合器捕捉跨网络信息流动。
- 提出S2B掩码机制,使感官刺激与血流信号对齐,实现严格因果建模。
- 在长时序预测中性能超越基线36%,适合神经科学与脑机接口研究者。
自然主义功能磁共振成像将大脑视为由连续感官输入驱动的动态预测引擎。然而,多模态输入与皮层网络复杂拓扑之间的时间尺度不匹配,阻碍了真实神经模拟中的因果前向演化建模。为此,我们提出BrainVista,一种多模态自回归框架,用于建模脑状态的因果演化。BrainVista引入网络级分词器以解耦系统特异性动态,并采用空间混合器头捕捉网络间信息流,同时保持功能边界。此外,我们提出新颖的刺激到脑(S2B)掩码机制,同步高频感官刺激与血流滤波信号,实现严格的仅历史因果条件。我们在Algonauts 2025、CineBrain和HAD数据集上验证该框架,达到最先进的fMRI编码性能。在长时序滚动设置下,模型相对于最强基线Algonauts 2025和CineBrain,模式相关性分别提升36.0%和33.3%。
原文摘要 · Abstract (English)
Naturalistic fMRI characterizes the brain as a dynamic predictive engine driven by continuous sensory streams. However, modeling the causal forward evolution in realistic neural simulation is impeded by the timescale mismatch between multimodal inputs and the complex topology of cortical networks. To address these challenges, we introduce BrainVista, a multimodal autoregressive framework designed to model the causal evolution of brain states. BrainVista incorporates Network-wise Tokenizers to disentangle system-specific dynamics and a Spatial Mixer Head that captures inter-network information flow without compromising functional boundaries. Furthermore, we propose a novel Stimulus-to-Brain (S2B) masking mechanism to synchronize high-frequency sensory stimuli with hemodynamically filtered signals, enabling strict, history-only causal conditioning. We validate our framework on Algonauts 2025, CineBrain, and HAD, achieving state-of-the-art fMRI encoding performance. In long-horizon rollout settings, our model yields substantial improvements over baselines, increasing pattern correlation by 36.0\% and 33.3\% on relative to the strongest baseline Algonauts 2025 and CineBrain, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。