用大模型提升印尼低资源语言翻译质量,效果显著。
NusaMT-7B: Machine Translation for Low-Resource Indonesian Languages with Large Language Models
- 基于LLaMA2-7B持续预训练+自学习+数据清洗,提升低资源语言翻译
- 在FLORES-200上对巴厘语、明加克语翻译提升最高达+6.69 spBLEU
- 适合关注语言保护与小语种AI的科研与公益项目
大型语言模型(LLMs)在高资源语言翻译中表现优异,但在低资源语言上受限于双语和单语语料稀缺及噪声问题,导致对齐困难,性能落后于最先进神经机器翻译(NMT)模型。本文提出NusaMT-7B,一个面向印尼低资源语言的LLM翻译模型,以巴厘语和明加克语为例。基于预训练的LLaMA2-7B,采用单语数据持续预训练、监督微调(SFT)、自学习和基于LLM的数据清洗方法,有效降低平行句对中的噪声。在FLORES-200多语言翻译基准测试中,NusaMT-7B在向巴厘语和明加克语翻译的spBLEU指标上最高提升达+6.69,但向高资源语言翻译时最差下降-3.38。结果表明,经微调的LLM可显著提升低资源语言翻译质量,助力语言保存与跨文化交流。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated exceptional promise in translation tasks for high-resource languages. However, their performance in low-resource languages is limited by the scarcity of both parallel and monolingual corpora, as well as the presence of noise. Consequently, such LLMs suffer with alignment and have lagged behind State-of-The-Art (SoTA) neural machine translation (NMT) models in these settings. This paper introduces NusaMT-7B, an LLM-based machine translation model for low-resource Indonesian languages, starting with Balinese and Minangkabau. Leveraging the pretrained LLaMA2-7B, our approach integrates continued pre-training on monolingual data, Supervised Fine-Tuning (SFT), self-learning, and an LLM-based data cleaner to reduce noise in parallel sentences. In the FLORES-200 multilingual translation benchmark, NusaMT-7B outperforms SoTA models in the spBLEU metric by up to +6.69 spBLEU in translations into Balinese and Minangkabau, but underperforms by up to -3.38 spBLEU in translations into higher-resource languages. Our results show that fine-tuned LLMs can enhance translation quality for low-resource languages, aiding in linguistic preservation and cross-cultural communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。