用词替换提升低资源语言的知识迁移效率
Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions

- 在预训练数据中对英文词汇进行词级替换,借助双语词表实现跨语言知识注入
- 在8种语言上显著提升下游任务表现,训练速度最快提升2倍
- 无需额外训练,仅需低成本双语词表,适合低资源语言场景
跨语言知识迁移对构建低资源语言的高性能多语言模型至关重要。当目标语言数据稀缺时,科学推理、常识推断和世界知识等下游任务所需知识主要依赖高资源语言获取,因此高效的知识迁移尤为关键。现有方法大多需要大量平行数据、翻译系统、辅助模型或额外训练阶段,这些在多数语言中难以获得。本文提出LINK——一种基于数据层面的干预方法,在预训练阶段通过双语词表对高资源语言(英语)训练数据中的部分词汇进行词级替换,从而提升知识迁移效果。该方法无需额外模型训练,仅需双语词表,而这类词表几乎可零成本获取。在五种模型规模下对八种语言的评估显示,目标语言下游任务性能显著提升,训练速度最快可达原方法的2倍。
原文摘要 · Abstract (English)
Cross-lingual knowledge transfer is critical for building high-performing multilingual language models for languages with insufficient training data. When target language data is scarce, the knowledge required for many downstream tasks involving scientific reasoning, commonsense inference, and world knowledge must be acquired primarily from the high-resource language, making effective knowledge transfer essential. Existing methods for improving such cross-lingual knowledge transfer require large amounts of parallel data, translation systems, auxiliary models, or additional training stages that are largely unavailable for many languages. We propose LINK - a data-level intervention method that improves knowledge transfer during model pretraining through lexical substitutions in high-resource part of pretraining data using bilingual vocabularies. For a given replacement ratio, randomly selected words in a portion of the high-resource (English) training corpus are swapped with their word-level translations, requiring no additional model training and only a bilingual vocabulary, which can be obtained at near-zero cost for virtually any language. Evaluation on eight languages across five model sizes shows notable improvements on downstream tasks in the target language, with up to a 2x speedup in training to reach equivalent performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。