用稀疏自编码器把深度模型与人脑视觉响应直接对齐,揭示信息处理机制。
Sparse Autoencoders Bridge The Deep Learning Model and The Brain
- 通过稀疏自编码器将模型激活与大脑fMRI信号匹配,实现端到端对齐。
- 最高相似度达0.76,证明SAE单元与皮层信号存在强对应关系。
- 适用于研究模型可解释性及脑-机映射,尤其适合认知神经科学与AI交叉方向。
我们提出SAE-BrainMap框架,利用稀疏自编码器(SAEs)直接对齐深度学习视觉模型表示与体素级fMRI响应。首先,在模型激活上训练逐层SAEs,计算SAE单元激活与相同自然图像刺激下皮层fMRI信号的余弦相似度,发现显著激活对应性(最大相似度达0.76)。基于此对齐,通过最优分配将最相似的SAE特征赋予每个体素,构建体素字典,验证了SAE单元保留了预定义区域(ROIs)的功能结构并表现出一致的选择性。最后,建立模型层与人类腹侧视觉通路间的细粒度层次映射;通过将体素字典激活投影至皮层表面,可视化深度模型中视觉信息的动态转换过程。结果显示,ViT-B/16$_{CLIP}$在早期层利用低层信息生成高层语义信息,并在后期重构低维表征。该成果建立了无需下游任务的深度网络与人脑视觉皮层的直接桥梁,为模型可解释性提供新视角。
原文摘要 · Abstract (English)
We present SAE-BrainMap, a novel framework that directly aligns deep learning visual model representations with voxel-level fMRI responses using sparse autoencoders (SAEs). First, we train layer-wise SAEs on model activations and compute the correlations between SAE unit activations and cortical fMRI signals elicited by the same natural image stimuli with cosine similarity, revealing strong activation correspondence (maximum similarity up to 0.76). Depending on this alignment, we construct a voxel dictionary by optimally assigning the most similar SAE feature to each voxel, demonstrating that SAE units preserve the functional structure of predefined regions of interest (ROIs) and exhibit ROI-consistent selectivity. Finally, we establish fine-grained hierarchical mapping between model layers and the human ventral visual pathway, also by projecting voxel dictionary activations onto individual cortical surfaces, we visualize the dynamic transformation of the visual information in deep learning models. It is found that ViT-B/16$_{CLIP}$ tends to utilize low-level information to generate high-level semantic information in the early layers and reconstructs the low-dimension information later. Our results establish a direct, downstream-task-free bridge between deep neural networks and human visual cortex, offering new insights into model interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。