长视频上下文提升多模态模型与大脑的对齐,尤其在叙事理解中。
How does longer temporal context enhance multimodal narrative video processing in the brain?
- 用3-24秒不同长度视频片段,研究大脑如何随时间整合多模态信息。
- 长时序上下文使多模态大模型与高级脑区的神经活动对齐显著提升。
- 任务提示影响大脑响应模式,揭示了高层区域对上下文依赖的动态调制。
理解人类和人工智能系统如何处理复杂叙事视频,是神经科学与机器学习交叉领域的基本挑战。本研究探讨视频片段的时序上下文长度(3–24秒)及叙事任务提示如何影响自然观影过程中大脑与模型的对齐。基于参与者观看完整电影的fMRI数据,分析对叙事上下文敏感的大脑区域如何在不同时间尺度上动态表征信息,并与模型特征对齐。结果发现,增加片段时长显著提升多模态大语言模型(MLLMs)的大脑对齐度,而单模态视频模型几乎无改善。较短时间窗口对应感知与早期语言区域,更长时间窗口则与高阶整合区域更一致,这一趋势在MLLMs中呈现层-皮层层级映射。此外,四种叙事任务提示实验显示,任务特异性诱导大脑区域对齐模式变化,并在高阶区域引发片段级调制的上下文依赖性转变。本研究将长篇叙事电影定位为研究长时序上下文整合在长上下文MLLM中的理想测试平台,及其与大脑叙事理解反应的关系。
原文摘要 · Abstract (English)
Understanding how humans and artificial intelligence systems process complex narrative videos is a fundamental challenge at the intersection of neuroscience and machine learning. This study investigates how the temporal context length of video clips (3--24 s clips) and the narrative-task prompting shape brain-model alignment during naturalistic movie watching. Using fMRI recordings from participants viewing full-length movies, we examine how brain regions sensitive to narrative context dynamically represent information over varying timescales and how these neural patterns align with model-derived features. We find that increasing clip duration substantially improves brain alignment for multimodal large language models (MLLMs), whereas unimodal video models show little to no gain. Further, shorter temporal windows align with perceptual and early language regions, while longer windows preferentially align higher-order integrative regions, mirrored by a layer-to-cortex hierarchy in MLLMs. Finally, experiments with four narrative-task prompts show that they elicit task-specific, region-dependent brain alignment patterns and context-dependent shifts in clip-level tuning in higher-order regions. Our work positions long-form narrative movies as a principled testbed for studying long-timescale temporal integration in long-context MLLMs and its relationship to cortical responses during narrative comprehension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。