arXiv:2411.12986cs.CLcs.LG2024-11ACL被引 7

用英语数据提升低资源语言模型性能,无需改模型或训练目标

Training Bilingual LMs with Data Constraints in the Targeted Language

  • 用高质英语数据辅助训练低资源目标语言模型
  • 英语数据越丰富,目标语言性能提升越明显
  • 适合资源匮乏语言的模型训练,尤其对近缘语言有效

大规模语言模型依赖海量网络爬取数据,当前进展主要集中在英语,因其拥有高质量预训练数据。对于多数其他语言,此类高质量数据难以获取。本文研究在目标语言数据不足的情况下,如何通过引入高质量辅助语言(如英语)的数据来提升预训练模型性能。通过量化使用富数据辅助语言与目标语言训练的性能差距,探索翻译系统的作用,分析数据受限时模型扩展的局限性,并提出新的辅助语言数据上采样方法。结果表明,更强的辅助数据集能带来性能提升,且无需修改模型或训练目标;尤其当英语预训练数据更丰富时,其优势可延伸至数据有限的目标语言场景。

原文摘要 · Abstract (English)

Large language models are trained on massive scrapes of the web, as required by current scaling laws. Most progress is made for English, given its abundance of high-quality pretraining data. For most other languages, however, such high quality pretraining data is unavailable. In this work, we study how to boost pretrained model performance in a target language with insufficient pretraining data for training a high performing language model, by enlisting data from an auxiliary language for which high quality data is available. We study this by quantifying the performance gap between training with data in a data-rich auxiliary language compared with training in the target language, exploring the benefits of translation systems, studying the limitations of model scaling when data is limited in the target languages, and proposing new methods for upsampling data from the auxiliary language. Our results show that stronger auxiliary datasets result in performance gains without modification to the model or training objective for close languages, and, in particular, that performance gains due to the development of more information-rich English pretraining datasets can extend to targeted language settings with limited data.

多语言低资源数据增强迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。