通过重加权多语言数据,提升大模型对低资源语言的性能。
XDoGE: Multilingual Data Reweighting to Enhance Language Inclusivity in LLMs
- 用代理模型在跨语言环境下优化数据分布,实现多语言均衡训练。
- 在六种语言上实验,低资源语言性能提升显著,高资源语言未受损。
- 适合关注语言公平性、多语言模型优化的研究者与开发者。
当前大型语言模型主要基于少数高资源语言(如英语)的海量文本训练,导致中低资源语言性能受限。为此,我们提出XDoGE方法:首先在领域重加权的DoGE算法基础上扩展为多语言版本,训练小型代理模型以优化语言分布;其次根据所得语言权重重新采样数据,并在全尺寸模型上从头或持续预训练(CPT)进行训练。实验涵盖六种语言:英语、西班牙语(高资源),葡萄牙语、加泰罗尼亚语(中资源),加利西亚语、巴斯克语(低资源)。使用Salamandra-2b模型和IberoBench框架评估数据重复与欠采样的影响。最终发布IberianLLM-7B-Instruct模型,基于原生预训练并结合XDoGE权重通过CPT进一步优化,专注于伊比利亚语言与英语。
原文摘要 · Abstract (English)
Current large language models (LLMs) are trained on massive amounts of text data, primarily from a few dominant languages. Studies suggest that this over-reliance on high-resource languages, such as English, hampers LLM performance in mid- and low-resource languages. To mitigate this problem, we propose to (i) optimize the language distribution by training a small proxy model within a domain-reweighing DoGE algorithm that we extend to XDoGE for a multilingual setup, and (ii) rescale the data and train a full-size model with the established language weights either from scratch or within a continual pre-training phase (CPT). We target six languages possessing a variety of geographic and intra- and inter-language-family relations, namely, English and Spanish (high-resource), Portuguese and Catalan (mid-resource), Galician and Basque (low-resource). We experiment with Salamandra-2b, which is a promising model for these languages. We investigate the effects of substantial data repetition on minor languages and under-sampling on dominant languages using the IberoBench framework for quantitative evaluation. Finally, we release a new promising IberianLLM-7B-Instruct model centering on Iberian languages and English that we pretrained from scratch and further improved using CPT with the XDoGE weights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。