通过语言算术识别并操控大模型中的语言特异性神经元。
Language Arithmetics: Towards Systematic Language Neuron Identification and Manipulation
- 用LAPE方法定位深层网络中控制语言行为的神经元。
- 语言算术可精准关闭或开启指定语言,效果优于简单替换。
- 高资源语言和语系相近的语言更易操控,适合多语言任务优化。
大型语言模型具备强大的多语言能力,但其语言特异性处理的神经机制仍不明确。本研究分析了Llama-3.1-8B、Mistral-Nemo-12B以及Aya-Expanse-8B与32B在21种语言类型多样的模型中,识别出控制语言行为的特异性神经元。基于语言激活概率熵(LAPE)方法,发现这些神经元集中在深层,非拉丁文字表现出更高专一性;相关语言共享重叠神经元,反映内部语言相似性表征。通过语言算术——即系统性的激活加法与乘法操作,实现对模型语言行为的引导,可有效关闭非期望语言并激活目标语言,优于简单替换策略。该干预在五项多语言任务中表现良好:语言强制、翻译、问答、理解与自然语言推理。高资源语言操控成功率更高,语言类型相似性提升干预效果。此外,跨语言神经元调控能提升下游性能,并揭示当神经元逐步失活时的内部‘备用’语言选择机制。代码已公开于https://github.com/d-gurgurov/Language-Neurons-Manipulation。
原文摘要 · Abstract (English)
Large language models (LLMs) exhibit strong multilingual abilities, yet the neural mechanisms behind language-specific processing remain unclear. We analyze language-specific neurons in Llama-3.1-8B, Mistral-Nemo-12B, and Aya-Expanse-8B & 32B across 21 typologically diverse languages, identifying neurons that control language behavior. Using the Language Activation Probability Entropy (LAPE) method, we show that these neurons cluster in deeper layers, with non-Latin scripts showing greater specialization. Related languages share overlapping neurons, reflecting internal representations of linguistic proximity. Through language arithmetics, i.e. systematic activation addition and multiplication, we steer models to deactivate unwanted languages and activate desired ones, outperforming simpler replacement approaches. These interventions effectively guide behavior across five multilingual tasks: language forcing, translation, QA, comprehension, and NLI. Manipulation is more successful for high-resource languages, while typological similarity improves effectiveness. We also demonstrate that cross-lingual neuron steering enhances downstream performance and reveal internal "fallback" mechanisms for language selection when neurons are progressively deactivated. Our code is made publicly available at https://github.com/d-gurgurov/Language-Neurons-Manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。