arXiv:2507.17897q-bio.NCcs.CV2025-07被引 5

用多模态循环模型预测人看电影时的脑部反应,效果优于多数参赛者。

Multimodal Recurrent Ensembles for Predicting Brain Responses to Naturalistic Movies (Algonauts 2025)

  • 分模态双向RNN捕捉视听语义随时间变化,再融合进第二层递归网络
  • 在1000个皮层区域上平均相关系数达0.2094,单区域最高0.63
  • 适合脑科学与多模态建模研究者,为未来脑编码任务提供可扩展基准

准确预测自然刺激下分布式皮层响应需要整合视觉、听觉和语义信息的时间动态。我们提出一种分层多模态循环集成模型,将预训练的视频、音频和语言嵌入映射到四名受试者观看近80小时电影期间记录的fMRI时间序列。各模态专用的双向RNN编码时间动态,其隐状态融合后输入第二层递归层,轻量级个体特异性头部输出1000个皮层区域的响应。训练采用复合均方误差-相关性损失,并通过渐进式课程学习逐步转移重点从早期感觉区到晚期关联区。平均100个模型变体进一步提升鲁棒性。该系统在竞赛排行榜中位列第三,整体皮尔逊相关系数达到0.2094,是所有参与者中单区域峰值分数最高(平均0.63),尤其在最具挑战性的受试者5上表现显著提升。该方法为未来的多模态脑编码基准任务建立了一个简洁且可扩展的基线。

原文摘要 · Abstract (English)

Accurately predicting distributed cortical responses to naturalistic stimuli requires models that integrate visual, auditory and semantic information over time. We present a hierarchical multimodal recurrent ensemble that maps pretrained video, audio, and language embeddings to fMRI time series recorded while four subjects watched almost 80 hours of movies provided by the Algonauts 2025 challenge. Modality-specific bidirectional RNNs encode temporal dynamics; their hidden states are fused and passed to a second recurrent layer, and lightweight subject-specific heads output responses for 1000 cortical parcels. Training relies on a composite MSE-correlation loss and a curriculum that gradually shifts emphasis from early sensory to late association regions. Averaging 100 model variants further boosts robustness. The resulting system ranked third on the competition leaderboard, achieving an overall Pearson r = 0.2094 and the highest single-parcel peak score (mean r = 0.63) among all participants, with particularly strong gains for the most challenging subject (Subject 5). The approach establishes a simple, extensible baseline for future multimodal brain-encoding benchmarks.

脑编码多模态循环网络fMRI预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。