从语言神经元视角解析大模型多语言对齐机制
How Does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons Perspective
- 提出三类神经元分类法:语言特有、相关与通用
- 发现对齐后共享语义空间神经元显著增加
- 适合研究大模型多语言能力与对齐原理的学者
多语言对齐是提升大语言模型多语言能力的有效方法,能将高资源语言的能力迁移至低资源语言。现有研究揭示了语言特异性神经元的存在,但发现大量神经元同时活跃于多种语言却无法被准确归类。本文提出一种三类神经元分类方法:语言特有神经元、语言相关神经元和通用神经元,并设计对应识别算法。基于不同神经元分布特征,将模型多语言推理过程划分为四个阶段:多语言理解、共享语义空间推理、多语言输出空间转换和词汇空间输出。系统对比对齐前后模型中各类神经元变化,分析“自发多语言对齐”现象。研究提供了基于神经元层级的实证结果,深化了对多语言对齐机制的理解。
原文摘要 · Abstract (English)
Multilingual Alignment is an effective and representative paradigm to enhance LLMs' multilingual capabilities, which transfers the capabilities from the high-resource languages to the low-resource languages. Meanwhile, some research on language-specific neurons provides a new perspective to analyze and understand LLMs' mechanisms. However, we find that there are many neurons that are shared by multiple but not all languages and cannot be correctly classified. In this work, we propose a ternary classification methodology that categorizes neurons into three types, including language-specific neurons, language-related neurons, and general neurons. And we propose a corresponding identification algorithm to distinguish these different types of neurons. Furthermore, based on the distributional characteristics of different types of neurons, we divide the LLMs' internal process for multilingual inference into four parts: (1) multilingual understanding, (2) shared semantic space reasoning, (3) multilingual output space transformation, and (4) vocabulary space outputting. Additionally, we systematically analyze the models before and after alignment with a focus on different types of neurons. We also analyze the phenomenon of "Spontaneous Multilingual Alignment". Overall, our work conducts a comprehensive investigation based on different types of neurons, providing empirical results and valuable insights to better understand multilingual alignment and multilingual capabilities of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。