用自然影像测试指令微调多模态模型,发现其脑活动对齐更强。
Task-conditioned probing of instruction-tuned multimodal LLMs: Region-specific brain alignment patterns under naturalistic stimuli
- 通过任务指令嵌入,探测多模态模型在自然视频中的表征。
- 指令微调模型比非微调模型脑对齐高约9%,优于基线15%-20%。
- 模型响应随任务变化,适合研究人机认知映射机制。
近期体素级多模态脑编码研究显示,多模态大语言模型(MLLMs)相比单模态模型具有更高脑对齐度。最近研究发现,指令微调的多模态(IT)模型能生成与任务相关的表征,且与脑活动高度一致,但以往评估多集中于单模态刺激或非指令微调模型。我们通过自然主义电影观看(视频+音频)期间记录的fMRI信号,预测来自六种视频和两种音频指令微调多模态模型的表示。结果表明,指令微调的视频MLLMs脑对齐度显著高于上下文学习(ICL)模型(~9%)、非指令微调模型(~15%)及单模态基线(~20%)。跨视频/音频任务的语言引导探测揭示了任务特异性的模型表征,且不同脑区差异明显。此外,ICL模型呈现强语义组织(r=0.78),而IT模型与指令文本语义耦合弱(r=0.14),支持任务条件子空间与更高脑对齐的关联。这些发现表明任务指令与更强脑-模型对齐相关,为联合信息处理机制研究开辟新路径。代码已公开:https://github.com/subbareddy248/mllm_videos。
原文摘要 · Abstract (English)
Recent voxel-wise multimodal brain encoding studies have shown that multimodal large language models (MLLMs) exhibit a higher degree of brain alignment compared to unimodal models. More recently, instruction-tuned multimodal (IT) models have been shown to generate task-specific representations that align strongly with brain activity, yet most prior evaluations focus on unimodal stimuli or non-instruction-tuned models under multimodal stimuli. We still lack a clear understanding of whether instruction-tuning is associated with IT-MLLMs organizing their representations around functional task demands or if they simply reflect surface semantics. To address this, we estimate brain alignment by predicting fMRI responses recorded during naturalistic movie watching (video with audio) from MLLM representations. Using instruction-specific embeddings from six video and two audio IT-MLLMs, across 13 video task instructions, we find that instruction-tuned video MLLMs show higher brain alignment than in-context learning (ICL) multimodal models (~9%), non-instruction-tuned multimodal models (~15%), and unimodal baselines (~20%). Our evaluation of MLLMs across video and audio tasks, and language-guided probing produces distinct task-specific MLLM representations that vary across brain regions. We also find that ICL models show strong semantic organization (r=0.78), while IT models show weak coupling to instruction-text semantics (r=0.14), consistent with task-conditioned subspaces associated with higher brain alignment. These findings are consistent with an association between task-specific instructions and stronger brain-MLLM alignment, and open new avenues for mapping joint information processing in both systems. We make the code publicly available [https://github.com/subbareddy248/mllm_videos].
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。