用计算损伤法揭示多语言模型中共享与语言特异的脑映射机制
Computational Lesions in Multilingual Language Models Separate Shared and Language-specific Brain Alignment

- 通过零值化关键参数,模拟神经损伤以测试模型对多语言脑响应的预测能力
- 共享核心受损导致全脑编码相关性下降60.32%,而语言特异性损伤仅影响对应母语预测
- 为多语言神经机制研究提供因果分析框架,适合认知神经科学与AI交叉研究者
人类大脑如何支持多种语言是神经科学的基本问题,也是检验多语言人工智能的有效标准。神经影像已发现跨语言的语言响应脑区,但无法确定其处理机制是共享还是语言特异。本文使用六种多语言大模型作为可控系统,通过零值化在多种语言中均重要的小参数集(共享)或特定语言关键参数(语言特异),构建“计算损伤”。随后比较完整模型与受损模型在112名母语为英语、中文和法语的参与者听自然故事时(共100分钟)的fMRI响应预测表现。结果显示,损伤紧凑的共享核心使全脑编码相关性相对降低60.32%;而语言特异性损伤虽保持嵌入空间的跨语言分离,但仅显著削弱对应母语的脑响应可预测性。结果支持“共享主干+嵌入特化”的架构,并为多语言脑-模型对齐研究提供因果框架。
原文摘要 · Abstract (English)
How the brain supports language across different languages is a basic question in neuroscience and a useful test for multilingual artificial intelligence. Neuroimaging has identified language-responsive brain regions across languages, but it cannot by itself show whether the underlying processing is shared or language-specific. Here we use six multilingual large language models (LLMs) as controllable systems and create targeted ``computational lesions'' by zeroing small parameter sets that are important across languages or especially important for one language. We then compare intact and lesioned models in predicting functional magnetic resonance imaging (fMRI) responses during 100 minutes of naturalistic story listening in native English, Chinese and French (112 participants). Lesioning a compact shared core reduces whole-brain encoding correlation by 60.32% relative to intact models, whereas language-specific lesions preserve cross-language separation in embedding space but selectively weaken brain predictivity for the matched native language. These results support a shared backbone with embedded specializations and provide a causal framework for studying multilingual brain-model alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。