arXiv:2510.16870cs.CV2025-10

用脑活动数据揭示视觉语言模型的类脑分层机制。

Uncovering Brain-Like Hierarchical Patterns in Vision-Language Models through fMRI-Based Neural Encoding

  • 通过神经元级分析,对比模型与人脑神经活动
  • 发现模型与大脑在多个功能网络中共享表征机制
  • 不同架构的模型表现出不同的类脑激活模式

尽管类脑人工智能已展现出良好前景,但当前对人工神经网络(ANN)与人类大脑处理机制之间关联的理解仍有限:(1) 单模态ANN研究无法捕捉大脑固有的多模态处理能力;(2) 多模态ANN研究主要关注高层输出,忽视个体神经元的作用。为此,我们提出一种新型神经元级分析框架,通过人脑活动(fMRI)视角研究视觉语言模型(VLMs)的多模态信息处理机制。该方法结合精细的人工神经元(AN)分析与基于fMRI的体素编码,考察了两种架构不同的VLM:CLIP与METER。结果显示:(1) AN能有效预测多个功能网络(包括语言、视觉、注意和默认模式网络)中的生物神经元(BN)活动,表明存在共享表征机制;(2) AN与BN均表现出功能冗余,神经表征重叠,反映大脑容错与协作的信息处理方式;(3) AN呈现极性模式,与BN一致,相反激活的BN在模型各层呈现镜像激活趋势,体现神经信息处理的复杂性和双向性;(4) 不同架构影响类脑特性:CLIP的独立分支显示模态特异性,而METER的跨模态设计产生统一跨模态激活。这些结果为VLM在神经元层面存在类脑分层处理提供了有力证据。

原文摘要 · Abstract (English)

While brain-inspired artificial intelligence(AI) has demonstrated promising results, current understanding of the parallels between artificial neural networks (ANNs) and human brain processing remains limited: (1) unimodal ANN studies fail to capture the brain's inherent multimodal processing capabilities, and (2) multimodal ANN research primarily focuses on high-level model outputs, neglecting the crucial role of individual neurons. To address these limitations, we propose a novel neuron-level analysis framework that investigates the multimodal information processing mechanisms in vision-language models (VLMs) through the lens of human brain activity. Our approach uniquely combines fine-grained artificial neuron (AN) analysis with fMRI-based voxel encoding to examine two architecturally distinct VLMs: CLIP and METER. Our analysis reveals four key findings: (1) ANs successfully predict biological neurons (BNs) activities across multiple functional networks (including language, vision, attention, and default mode), demonstrating shared representational mechanisms; (2) Both ANs and BNs demonstrate functional redundancy through overlapping neural representations, mirroring the brain's fault-tolerant and collaborative information processing mechanisms; (3) ANs exhibit polarity patterns that parallel the BNs, with oppositely activated BNs showing mirrored activation trends across VLM layers, reflecting the complexity and bidirectional nature of neural information processing; (4) The architectures of CLIP and METER drive distinct BNs: CLIP's independent branches show modality-specific specialization, whereas METER's cross-modal design yields unified cross-modal activation, highlighting the architecture's influence on ANN brain-like properties. These results provide compelling evidence for brain-like hierarchical processing in VLMs at the neuronal level.

类脑计算视觉语言模型神经科学多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。