arXiv:2412.07298cs.CL2024-12ICLR被引 2

揭示代码大模型多语言能力演化机制,提出'巴别塔假说'

The Rise and Down of Babel Tower: Investigating the Evolution Process of Multilingual Code Large Language Model

  • 通过追踪模型内部状态,发现多语言能力从共用主语言知识逐步分化
  • 提出优化预训练语料的新方法,显著提升多语言代码模型性能
  • 适合研究多语言模型演化与数据构建的学者参考

大型语言模型(LLMs)展现出显著的多语言能力,但其在预训练过程中能力发展的机制尚不明确。本文以代码大模型为实验平台,探究了LLM在预训练过程中多语言能力的演化过程。基于观察,我们提出巴别塔假说,描述了模型获取新语言能力的完整过程:多种语言初期共享一个由主语言主导的知识系统,随后逐步发展出语言特异的知识系统。我们通过识别工作语言和语言迁移神经元,验证了该假说,实验结果表明模型内部状态变化与巴别塔假说一致。基于此,我们提出一种新型预训练语料构建方法,显著优于原始语料训练的模型。该假说为设计最优多语言能力的预训练数据分布提供了新视角。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown significant multilingual capabilities. However, the mechanisms underlying the development of these capabilities during pre-training are not well understood. In this paper, we use code LLMs as an experimental platform to explore the evolution of multilingual capabilities in LLMs during the pre-training process. Based on our observations, we propose the Babel Tower Hypothesis, which describes the entire process of LLMs acquiring new language capabilities. During the learning process, multiple languages initially share a single knowledge system dominated by the primary language and gradually develop language-specific knowledge systems. We then validate the above hypothesis by tracking the internal states of the LLMs through identifying working languages and language transferring neurons. Experimental results show that the internal state changes of the LLM are consistent with our Babel Tower Hypothesis. Building on these insights, we propose a novel method to construct an optimized pre-training corpus for multilingual code LLMs, which significantly outperforms LLMs trained on the original corpus. The proposed Babel Tower Hypothesis provides new insights into designing pre-training data distributions to achieve optimal multilingual capabilities in LLMs.

多语言模型代码生成预训练知识演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。