用视频音频文本融合模型预测大脑对影视的反应,效果优于以往方法。
VIBE: Video-Input Brain Encoder for fMRI Response Modeling
- 多模态特征经融合变压器编码,再通过带旋转位置编码的预测变压器解码
- 在内分布影片上达到0.3225的皮尔逊相关系数,外分布影片为0.2125
- 适用于脑功能研究、神经影像建模,尤其适合关注多模态输入的团队
我们提出VIBE,一种两阶段Transformer模型,通过融合多模态视频、音频和文本特征来预测功能性磁共振成像(fMRI)活动。来自开源模型(Qwen2.5、BEATs、Whisper、SlowFast、V-JEPA)的表征经由模态融合变压器整合,并通过带有旋转位置嵌入的预测变压器进行时序解码。该模型在CNeuroMod数据集的65小时电影数据上训练,并通过20个随机种子集成。在同分布的Friends S07上达到0.3225的皮尔逊相关系数,在六个异分布电影上为0.2125。该架构早期版本在Algonauts 2025挑战赛中分别获得第一阶段冠军和总排名第二。
原文摘要 · Abstract (English)
We present VIBE, a two-stage Transformer that fuses multi-modal video, audio, and text features to predict fMRI activity. Representations from open-source models (Qwen2.5, BEATs, Whisper, SlowFast, V-JEPA) are merged by a modality-fusion transformer and temporally decoded by a prediction transformer with rotary embeddings. Trained on 65 hours of movie data from the CNeuroMod dataset and ensembled across 20 seeds, VIBE attains mean parcel-wise Pearson correlations of 0.3225 on in-distribution Friends S07 and 0.2125 on six out-of-distribution films. An earlier iteration of the same architecture obtained 0.3198 and 0.2096, respectively, winning Phase-1 and placing second overall in the Algonauts 2025 Challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。