arXiv:2608.12717cs.LGcs.CL2026-08

用脑神经成像方法解析大模型内部功能分工,发现语音加工区域更敏感。

Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia

论文配图:Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia
图 1 · 摘自论文原文
  • 将人类脑损伤研究的减法分析移植到大模型,通过扰动分层定位功能区域
  • 在模型与中风患者中均复现语音优势分离现象,且结果可重复
  • 为模型功能特异性提供可验证的空间化测试框架,适合认知神经科学与AI交叉研究

大型语言模型的机制可解释性缺乏空间分辨、可证伪的工具来检验内部组件是否专用于特定认知操作。我们借鉴人类神经影像学的标准框架——减法分析,将其从生物大脑迁移至受扰动的Transformer模型,并在两类基底上并行应用相同逻辑。基于此前显示层扰动的LLaVA-1.6-Vicuna-13B错误模式与失语症患者病灶模式匹配的Brain-LLM Unified Model(BLUM),我们提出PRISM(Perturbation-based Regional Interpretability through Subtraction Mapping)。PRISM映射费城命名测试的七类临床类别,成对减去错误类别,将每个扰动种子视为一个受试者,在层轴上使用无阈值簇增强进行组分析。我们在213名慢性中风失语症患者上运行结构匹配分析,采用相关性差异病灶-症状映射,并在保留数据集上复现。设计在受试者维度(种子/患者)、空间维度(层/图谱化皮层)和阈值处理上一致,但对比算子不同:模型为组内错误比例差异,大脑为组间相关性差异。两者均恢复出稳健的语音优势分离现象,包括深层层簇和额颞叶皮层簇,且均可复现;语义优势方向则为一致但不显著的趋势。因此,PRISM为变压器语言模型的功能特异性主张提供了可证伪、空间分辨的测试方法。确认性区域干预(PRISM第三阶段)以确立最强因果机制声明,留待后续工作。

原文摘要 · Abstract (English)

Mechanistic interpretability of large language models lacks spatially resolved, falsifiable tools for testing whether internal components are specialized for distinct cognitive operations. We adapt subtraction analysis, the standard framework of human neuroimaging, from biological brains to perturbed transformers, and apply the same logic to both substrates in parallel. Building on the Brain-LLM Unified Model (BLUM), which showed that layer-perturbed LLaVA-1.6-Vicuna-13B error profiles match the lesion patterns of aphasic patients, we develop PRISM (Perturbation-based Regional Interpretability through Subtraction Mapping). PRISM maps the seven clinical Philadelphia Naming Test categories, subtracts error classes pairwise, and treats each perturbation seed as a subject in a group analysis with threshold-free cluster enhancement along the layer axis. We run a structurally matched analysis on 213 chronic post-stroke aphasia patients using correlation-difference lesion-symptom mapping, and replicate both sides on held-out splits. The designs match in subject dimension (seeds, patients), spatial dimension (layers, atlas-parcellated cortex) and thresholding, but the contrast operator differs: a within-subject error-proportion difference for the LLM, a between-subject correlation difference for the cortex. Both substrates recover a robust phonemic-favoring dissociation, a deep layer cluster and a frontal-perisylvian cortical cluster, both replicating; the semantic-favoring direction is a consistently signed but non-significant trend on both. PRISM thus gives a falsifiable, spatially resolved test of functional-specialization claims in transformer language models. A confirmatory ROI-level intervention (PRISM Stage 3) licensing the strongest causal-mechanism claim is left to subsequent work.

可解释性语言模型神经对照大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。