中文大模型在脑-模型对齐中表现更优,非英语天生占优。
Brain-LLM Alignment Tracks Training Data, Not Typology
- 用跨语言fMRI数据验证模型对齐,发现训练语言主导对齐模式
- 中文主导模型在中文大脑中对齐最佳,英语对齐反而最差
- 语法差异是影响对齐的核心因素,尤其影响额下回区域
脑-大模型对齐在英语中已得到充分验证,但语言网络在神经解剖上具有跨语言普遍性。这种对齐是否能跨语言泛化?其变化由什么决定?我们基于112名参与者(英语、汉语、法语)的fMRI数据,使用7个涵盖英语主导、中文主导及多语言架构的大模型进行测试。核心发现:训练语言主导性而非英语固有属性决定了对齐模式——中文主导模型Baichuan2-7B(与LLaMA-2-7B同架构)完全逆转梯度,对中文大脑对齐最佳,对英语大脑对齐最差。除训练主导性外,形式类型学距离独立影响对齐退化;语法相关脑区(IFG)的梯度比词义区域(PTL)陡峭2.3倍,且分词丰富度解释了约60%跨语言最优编码层的偏移。结果表明,脑-模型对齐中的‘英语优势’实为训练数据构成的产物,剩余变异则反映了集中于句法处理的真实类型结构。
原文摘要 · Abstract (English)
Brain-LLM alignment is well established in English, yet the brain's language network is neuroanatomically universal across languages. Does alignment also generalize cross-linguistically, and what governs the variation? We test this using fMRI data from 112 participants across English, Chinese, and French (the Le Petit Prince corpus) and seven LLMs spanning English-dominant, Chinese-dominant, and multilingual architectures. Our central finding is that training-language dominance, not an inherent property of English, drives the alignment pattern: a Chinese-dominant model (Baichuan2-7B), architecture-matched to LLaMA-2-7B, reverses the gradient entirely, aligning best with Chinese brains and worst with English. Beyond training dominance, formal typological distance independently covaries with alignment degradation, syntax-associated brain regions (IFG) show $2.3\times$ steeper typological gradients than lexico-semantic regions (PTL), and tokenization fertility accounts for $\sim$60% of a cross-linguistic shift in optimal encoding layer. These results reveal that the apparent "English advantage" in brain-LLM alignment is an artifact of training data composition, while the remaining variation reflects genuine typological structure concentrated in syntactic processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。