发现多语言大模型中负责跨语言表征迁移的关键神经元。
The Transfer Neurons Hypothesis: An Underlying Mechanism for Language Latent Space Transitions in Multilingual LLMs
- 识别出MLP模块中专门负责语言间表征转移的神经元。
- 这些转移神经元对多语言推理任务至关重要,影响模型表现。
- 揭示了语言特异神经元在跨空间移动中的辅助作用。
近期研究提出解码器类大语言模型处理多语言输入的框架:早期层将输入转化为以英语为中心且语言无关的表征;中间层在以英语为中心的潜在空间内进行推理;最终层将这些表征转换回语言特定的潜在空间以生成输出。然而,这种转换的内部动态及底层机制仍不明确。为深入理解该框架,本文提出并实证验证了‘转移神经元假说’:某些MLP模块中的神经元负责在语言特定潜在空间与共享语义潜在空间之间传递表征。此外,我们表明,近期研究识别的语言特异性神经元的一个功能是促进潜在空间间的移动。最后,我们证明转移神经元对多语言大模型的推理至关重要。
原文摘要 · Abstract (English)
Recent studies have suggested a processing framework for multilingual inputs in decoder-based LLMs: early layers convert inputs into English-centric and language-agnostic representations; middle layers perform reasoning within an English-centric latent space; and final layers generate outputs by transforming these representations back into language-specific latent spaces. However, the internal dynamics of such transformation and the underlying mechanism remain underexplored. Towards a deeper understanding of this framework, we propose and empirically validate The Transfer Neurons Hypothesis: certain neurons in the MLP module are responsible for transferring representations between language-specific latent spaces and a shared semantic latent space. Furthermore, we show that one function of language-specific neurons, as identified in recent studies, is to facilitate movement between latent spaces. Finally, we show that transfer neurons are critical for reasoning in multilingual LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。