arXiv:2502.14671cs.CLcs.AI2025-02被引 6

用AI解释方法揭示大模型如何映射大脑语言处理机制

Explanations of Large Language Models Explain Language Representations in the Brain

  • 用梯度归因法量化每个词对模型预测的贡献
  • 归因结果在听觉皮层显著匹配脑活动,优于模型内部表示
  • 早期层关注词性,晚期层侧重位置信息,反映不同脑区对应

大语言模型(LLM)表征与语言加工时的大脑活动存在对齐,但其成因尚不明确。本文测试可解释AI(XAI)能否解答此问题:通过归因方法量化输入词语对模型下一个词预测的贡献,并用这些解释预测参与者听叙事时的fMRI数据。结果显示,基于梯度的归因方法在脑活动预测中具有稳健对齐性,能解释独立于声学与词频混杂因素的变异,在早期听觉区域的表现优于模型内部表征。进一步利用导纳(conductance)将归因扩展至各网络层,发现早期层对词性敏感,更匹配听觉皮层;最终层的归因则以位置信息为主导,表现出广泛的皮层对齐。研究证明,基于归因的解释不仅能衡量模型-脑对齐,还能揭示其背后的认知含义。

原文摘要 · Abstract (English)

Large Language Model (LLM) representations are known to align with brain activity during language processing, but it remains unclear what drives this alignment. We test whether explainable AI (XAI) can help answer this: using attribution methods, we quantify the contribution of each input word to an LLM's next-word predictions and use these explanations to predict fMRI data from participants listening to narratives. We find that gradient-based attribution methods robustly align with brain activity, contribute unique variance beyond acoustic and word-rate confounds, and outperform internal representations in early auditory regions. Using conductance, we extend attribution from words to individual layers, asking what each layer's attribution reveals about the model's computation and how this relates to its brain alignment. Early layers show greater word-type sensitivity and align preferentially with auditory regions, whereas the final layer's attribution is dominated by positional information and exhibits broad cortical alignment. Together, these findings demonstrate that attribution-based explanations can be used not only to measure LLM--brain alignment but to characterize what it reflects.

大模型解释脑机对齐可解释AIfMRI分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。