用中风患者脑损伤数据验证大模型语言能力,发现错误模式与人类脑区损伤高度对应。
Stroke Lesions as a Rosetta Stone for Language Model Interpretability
- 以中风患者脑损伤-症状映射为外部标准,评估大模型扰动后的语言错误
- 大模型错误模式在67%图片命名和68.3%句式补全中预测出真实脑损伤位置
- 语义错误对应腹侧通路损伤,发音错误对应背侧通路损伤,提示计算机制相似
大语言模型(LLMs)表现出强大语言能力,但缺乏验证其组件必要性的方法。现有可解释性研究依赖内部指标,缺乏外部验证。本文提出脑-语言模型统一框架(BLUM),利用百年来公认的脑-行为因果关系金标准——病灶-症状映射,作为评估大模型扰动效应的外部参照。基于410名慢性中风失语症患者的临床数据,训练了从行为错误模式预测脑损伤位置的模型,对Transformer各层进行系统性扰动,将扰动后的大模型与人类患者进行相同临床评估,并将模型错误模式投影到人类病灶空间。结果显示,大模型错误模式与人类错误模式高度相似:在67%的图片命名任务中,预测病灶位置显著优于随机水平(p < 10⁻²³);在68.3%的句子补全任务中同样显著(p < 10⁻⁶¹)。语义主导错误对应腹侧通路损伤模式,发音主导错误对应背侧通路损伤模式。该成果开辟了大模型可解释性新路径,使临床神经科学成为评估人工语言系统的外部验证依据,也提示行为一致性可能反映共享的计算原理。
原文摘要 · Abstract (English)
Large language models (LLMs) have achieved remarkable capabilities, yet methods to verify which model components are truly necessary for language function remain limited. Current interpretability approaches rely on internal metrics and lack external validation. Here we present the Brain-LLM Unified Model (BLUM), a framework that leverages lesion-symptom mapping, the gold standard for establishing causal brain-behavior relationships for over a century, as an external reference structure for evaluating LLM perturbation effects. Using data from individuals with chronic post-stroke aphasia (N = 410), we trained symptom-to-lesion models that predict brain damage location from behavioral error profiles, applied systematic perturbations to transformer layers, administered identical clinical assessments to perturbed LLMs and human patients, and projected LLM error profiles into human lesion space. LLM error profiles were sufficiently similar to human error profiles that predicted lesions corresponded to actual lesions in error-matched humans above chance in 67% of picture naming conditions (p < 10^{-23}) and 68.3% of sentence completion conditions (p < 10^{-61}), with semantic-dominant errors mapping onto ventral-stream lesion patterns and phonemic-dominant errors onto dorsal-stream patterns. These findings open a new methodological avenue for LLM interpretability in which clinical neuroscience provides external validation, establishing human lesion-symptom mapping as a reference framework for evaluating artificial language systems and motivating direct investigation of whether behavioral alignment reflects shared computational principles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。