融合多模型生成的电影脑响应预测,提升视觉与语言理解能力。
Stacked Regression using Off-the-shelf, Stimulus-tuned and Fine-tuned Neural Networks for Predicting fMRI Brain Responses to Movies (Algonauts 2025 Report)
- 用堆叠回归整合多种预训练模型的输出
- 结合字幕与摘要增强文本输入,显著提升预测精度
- 适合关注多模态脑科学建模的研究者
我们提交了参加 Algonauts 2025 挑战赛的方案,目标是预测人脑对电影刺激的 fMRI 响应。方法融合了大语言模型、视频编码器、音频模型及视觉语言模型的多模态表征,包含即插即用和微调版本。通过引入详细字幕和摘要增强文本输入,并探索语言与视觉模型的刺激调优与微调策略,提升了模型表现。各模型预测结果通过堆叠回归进行融合,取得良好效果。以团队名 Seinfeld 提交,排名第十。所有代码与资源均已公开,助力多模态脑活动建模研究。
原文摘要 · Abstract (English)
We present our submission to the Algonauts 2025 Challenge, where the goal is to predict fMRI brain responses to movie stimuli. Our approach integrates multimodal representations from large language models, video encoders, audio models, and vision-language models, combining both off-the-shelf and fine-tuned variants. To improve performance, we enhanced textual inputs with detailed transcripts and summaries, and we explored stimulus-tuning and fine-tuning strategies for language and vision models. Predictions from individual models were combined using stacked regression, yielding solid results. Our submission, under the team name Seinfeld, ranked 10th. We make all code and resources publicly available, contributing to ongoing efforts in developing multimodal encoding models for brain activity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。