用合成数据继续预训练,让多语言大模型更好支持低资源语言
Adapting Multilingual LLMs to Low-Resource Languages using Continued Pre-training and Synthetic Corpus
- 用真实+翻译生成的合成语料持续预训练模型
- 在4000亿词上训练后,印地语任务表现达最优
- 适合想提升低资源语言能力的研究者与开发者
多语言大模型虽支持多种语言,但在低资源语言上表现不佳。本文强调持续预训练与基于翻译的合成语料对提升低资源语言性能的重要性。以印地语为例,我们推出了基于Nemotron-Mini 4B的双语小模型Nemotron-Mini-Hindi 4B,使用真实与合成的印地语+英语混合 tokens,在4000亿词上进行持续预训练。实验表明,该模型在印地语基准测试中达到当前最优表现,同时在英语任务上保持竞争力。此外,持续预训练显著提升了模型整体事实准确性。消融实验显示,仅靠对齐无法实现印地语对话能力与事实准确性的提升,必须依赖印地语预训练。
原文摘要 · Abstract (English)
Multilingual LLMs support a variety of languages; however, their performance is suboptimal for low-resource languages. In this work, we emphasize the importance of continued pre-training of multilingual LLMs and the use of translation-based synthetic pre-training corpora for improving LLMs in low-resource languages. We conduct our study in the context of the low-resource Indic language Hindi. We introduce Nemotron-Mini-Hindi 4B, a bilingual SLM supporting both Hindi and English, based on Nemotron-Mini 4B. The model is trained using a mix of real and synthetic Hindi + English tokens, with continuous pre-training performed on 400B tokens. We demonstrate that both the base and instruct models achieve state-of-the-art results on Hindi benchmarks while remaining competitive on English tasks. Additionally, we observe that the continued pre-training approach enhances the model's overall factual accuracy. We perform an ablation study to highlight the impact of Hindi pre-training, showing significant improvements in Hindi chat capabilities and factual accuracy, which cannot be achieved through Hindi alignment alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。