为4种低资源印地语族语言构建高效翻译模型,提升小语种机器翻译性能。
SPRING Lab IITM's submission to Low Resource Indic Language Translation Shared Task
- 整合多源数据并用回译扩充语料,缓解双语数据稀缺问题。
- 在3种语言上微调NLLB 3.3B模型,1种语言通过特殊标记训练新模型。
- 适用于小语种翻译研究者及资源匮乏语言的本土化应用开发。
我们为四种低资源印地语族语言——Khasi、Mizo、Manipuri和Assamese——构建了鲁棒的翻译模型。方法涵盖从数据收集与预处理到训练与评估的完整流程,利用WMT任务数据集、BPCC、PMIndia和OpenLanguageData等多源数据。针对双语数据匮乏问题,对Mizo和Khasi使用单语数据进行回译,显著扩大训练语料。在Assamese、Mizo和Manipuri上微调预训练的NLLB 3.3B模型,性能优于基线。对于NLLB不支持的Khasi语言,引入特殊标记并在自建语料上训练。训练过程包含掩码语言建模,随后进行英-印及印-英翻译的微调。
原文摘要 · Abstract (English)
We develop a robust translation model for four low-resource Indic languages: Khasi, Mizo, Manipuri, and Assamese. Our approach includes a comprehensive pipeline from data collection and preprocessing to training and evaluation, leveraging data from WMT task datasets, BPCC, PMIndia, and OpenLanguageData. To address the scarcity of bilingual data, we use back-translation techniques on monolingual datasets for Mizo and Khasi, significantly expanding our training corpus. We fine-tune the pre-trained NLLB 3.3B model for Assamese, Mizo, and Manipuri, achieving improved performance over the baseline. For Khasi, which is not supported by the NLLB model, we introduce special tokens and train the model on our Khasi corpus. Our training involves masked language modelling, followed by fine-tuning for English-to-Indic and Indic-to-English translations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。