arXiv:2606.06857cs.CL2026-06

用稀疏编码特征解析大脑对语言的响应,发现脑区偏好通用语言信息。

Interpreting Brain Responses to Language with Sparse Features from Language Models

论文配图:Interpreting Brain Responses to Language with Sparse Features from Language Models
图 1 · 摘自论文原文
  • 用分层稀疏自编码器提取语言模型的稀疏特征,结合意外度预测
  • 识别出大脑对人物相关内容敏感的神经元群,此前未被描述
  • 脑区响应与模型中最具泛化性的特征最匹配,非任意特征

认知神经科学的核心目标之一是刻画人类语言皮层所表征的特征。人工语言模型(LM)已成为解决该问题的强大工具,但将生物与人工表示关联的研究常被批评为将一个黑箱与另一个黑箱对接。本文提出增强型稀疏编码模型,用分层组织的稀疏自编码器(SAE)特征替代密集的LM隐藏状态,并显式包含意外度作为预测因子。利用此方法,我们(i)解释了神经响应的机制,(ii)检验了模型-脑对齐是否反映模型表征中的主要或个性变异。基于8名参与者在7T fMRI下聆听200个语言多样句子的数据集,我们首先验证了建模框架:成功复现了以往关于处理难度和语义抽象性敏感的体素群的解释。随后,我们解释了一个此前未被描述但可靠的体素群,发现其对人物相关内容敏感。接着,我们表明额颞语言网络各区域由一组共同特征预测,但前部区域即使无LM特征,仅靠意外度也能较好解释。最后,我们证明大脑语言响应并非可由任意一组LM特征预测,而是最好由能捕捉到模型表征中最大共通信息的特征解释,表明大脑与语言模型之间存在非平凡的对应关系。

原文摘要 · Abstract (English)

A central goal of cognitive neuroscience is to characterize the features that are represented by human language cortex. Artificial language models (LMs) have emerged as a powerful tool to address this challenge, but studies relating biological and artificial representations are often criticized as relating one black box to another. The present work introduces Augmented Sparse Encoding Models, an encoding framework that replaces dense LM hidden states with hierarchically-organized sparse autoencoder (SAE) features, while explicitly including surprisal as a predictor. Using this approach, we (i) produce interpretations of neural responses and (ii) test whether model-brain alignment reflects primary or idiosyncratic variation in LM representations. Using a high-field 7T fMRI dataset of eight participants listening to 200 linguistically diverse sentences, we first validate our modeling framework by recovering previous interpretations of voxel populations tuned to processing difficulty and meaning abstractness. We then interpret a previously-uncharacterized (but reliable) voxel population and find that it is tuned to people-related content. Next, we show that the fronto-temporal human language network is predicted by a common set of features across its constituent regions, but find that frontal regions are relatively well-explained by surprisal alone, even in the absence of LM-based features. Finally, we show that brain responses during language processing are not merely predictable from an arbitrary set of LM features. Rather, brain responses are best explained by the features that tend to capture the most general information encoded in LM representations, suggesting a nontrivial correspondence between brain and LM language representation.

脑机接口语言模型神经科学稀疏编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。