通过合成语码转换提升多语言模型的跨语言对齐能力
Investigating and Scaling up Code-Switching for Multilingual Language Model Pre-Training
- 分析预训练语料中的四种语码转换类型及其分布
- 合成语码转换数据显著提升多语言性能,尤其在低资源语言上
- 适合研究多语言模型训练与语言对齐的学者
大型语言模型虽在预训练数据中存在极端语言不平衡,仍展现出出色的多语言能力。本文深入分析预训练语料,发现语码转换(即上下文中语言交替)是关键因素。我们识别出四种语码转换类型,并将其归入两个象限,发现其比例不均且对语言迁移影响各异。为探索语码转换在预训练中对语言对齐的作用,我们采用合成语码转换策略,持续扩大其数据量。实验表明,该方法在多个基准测试和表征空间中均带来显著提升,能有效增强高、中、低资源语言的泛化能力,且适用于不同质量的预训练语料。
原文摘要 · Abstract (English)
Large language models (LLMs) exhibit remarkable multilingual capabilities despite the extreme language imbalance in the pre-training data. In this paper, we closely examine the reasons behind this phenomenon, focusing on the pre-training corpus. We find that the existence of code-switching, alternating between different languages within a context, is key to multilingual capabilities. We conduct an analysis to investigate code-switching in the pre-training corpus, examining its presence and categorizing it into four types within two quadrants. We then assess its impact on multilingual performance. These types of code-switching data are unbalanced in proportions and demonstrate different effects on facilitating language transfer. To better explore the power of code-switching for language alignment during pre-training, we investigate the strategy of synthetic code-switching. We continuously scale up the synthetic code-switching data and observe remarkable improvements in both benchmarks and representation space. Extensive experiments indicate that incorporating synthetic code-switching data enables better language alignment and generalizes well to high, medium, and low-resource languages with pre-training corpora of varying qualities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。