arXiv:2505.20029q-bio.NCcs.AI2025-05ICLR被引 8

多模态大模型能更好模拟大脑视觉处理,尤其在指令引导下。

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)

  • 用指令微调的多模态模型生成文本嵌入,可更准确预测大脑视觉活动。
  • 10种指令测试中,模型脑对齐效果优于纯视觉模型,接近CLIP水平。
  • 模型能捕捉计数与识别等任务相关概念,适合研究神经编码机制者阅读。

基于Transformer的语言模型虽未显式训练以模仿脑电记录,却展现出与大脑活动的惊人一致性。模型规模增大、指令微调和多模态化进展提升了其表征与神经数据的对齐程度。近期出现的新一代指令微调多模态大语言模型(MLLMs)在开放式多模态视觉任务中表现出卓越的零样本能力。然而,当使用自然指令提示时,这些模型是否带来更好的脑对齐,并有效捕捉指令特异性表示尚不明确。为此,我们首次研究了脑对齐问题:通过测量不同指令下MLLM输出嵌入对观看自然场景时神经视觉活动的可预测性。10种不同指令的实验表明,相较于仅视觉模型,MLLMs显著提升脑对齐表现,且性能可媲美非指令微调的多模态模型如CLIP。此外发现,尽管这些模型能生成高质量的任务特定响应,但并非所有指令都促进脑对齐。通过调整指令,我们使模型编码与输入图像相关的指令特异性视觉概念,证明其能有效捕捉计数与识别相关概念,表现出与大脑活动的高度一致性。值得注意的是,多数大脑编码模型的解释方差在图像描述与其他指令的MLLM嵌入间共享。结果表明,增强MLLM对任务信息的捕捉能力,有助于更好区分各类指令,从而提升其预测脑响应的精度。

原文摘要 · Abstract (English)

Transformer-based language models, though not explicitly trained to mimic brain recordings, have demonstrated surprising alignment with brain activity. Progress in these models-through increased size, instruction-tuning, and multimodality-has led to better representational alignment with neural data. Recently, a new class of instruction-tuned multimodal LLMs (MLLMs) have emerged, showing remarkable zero-shot capabilities in open-ended multimodal vision tasks. However, it is unknown whether MLLMs, when prompted with natural instructions, lead to better brain alignment and effectively capture instruction-specific representations. To address this, we first investigate brain alignment, i.e., measuring the degree of predictivity of neural visual activity using text output response embeddings from MLLMs as participants engage in watching natural scenes. Experiments with 10 different instructions show that MLLMs exhibit significantly better brain alignment than vision-only models and perform comparably to non-instruction-tuned multimodal models like CLIP. We also find that while these MLLMs are effective at generating high-quality responses suitable to the task-specific instructions, not all instructions are relevant for brain alignment. Further, by varying instructions, we make the MLLMs encode instruction-specific visual concepts related to the input image. This analysis shows that MLLMs effectively capture count-related and recognition-related concepts, demonstrating strong alignment with brain activity. Notably, the majority of the explained variance of the brain encoding models is shared between MLLM embeddings of image captioning and other instructions. These results suggest that enhancing MLLMs' ability to capture task-specific information could lead to better differentiation between various types of instructions, and thereby improving their precision in predicting brain responses.

多模态模型脑机接口指令微调神经编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。