22种印地语系语言的多语言翻译模型,性能接近474M参数大模型。
NLIP_Lab-IITH Multilingual MT System for WAT24 MT Shared Task
- 用对齐一致性目标预训练印地语系语言,结合双语词典替换源句词汇。
- 基于高质量小样本数据微调方向特异性多语言模型,实现243M参数规模。
- 在英-印地语和印地语-英双向任务中表现优异,超越多数轻量级模型。
本文介绍了NLIP Lab为WAT24多语言印地语系机器翻译共享任务构建的多语言翻译系统,涵盖4个语系共22种计划语言。我们采用对齐一致性目标对印地语系语言进行预训练,并利用双语词典替换源句子中的词汇。此外,通过少量高质量种子数据对语言方向特异性多语言模型进行微调。主提交模型为包含22种印地语系语言的243M参数多语言翻译模型。在IN22-Gen基准测试中,英→印地语方向平均chrF++得分为46.80,BLEU为18.19;印地语→英方向平均chrF++为56.34,BLEU为30.82。在In22-Conv基准中,英→印地语方向chrF++为43.43,BLEU为16.58;印地语→英方向分别为52.44和29.77。该模型性能与474M参数的IndicTransv1模型相当。
原文摘要 · Abstract (English)
This paper describes NLIP Lab's multilingual machine translation system for the WAT24 shared task on multilingual Indic MT task for 22 scheduled languages belonging to 4 language families. We explore pre-training for Indic languages using alignment agreement objectives. We utilize bi-lingual dictionaries to substitute words from source sentences. Furthermore, we fine-tuned language direction-specific multilingual translation models using small and high-quality seed data. Our primary submission is a 243M parameters multilingual translation model covering 22 Indic languages. In the IN22-Gen benchmark, we achieved an average chrF++ score of 46.80 and 18.19 BLEU score for the En-Indic direction. In the Indic-En direction, we achieved an average chrF++ score of 56.34 and 30.82 BLEU score. In the In22-Conv benchmark, we achieved an average chrF++ score of 43.43 and BLEU score of 16.58 in the En-Indic direction, and in the Indic-En direction, we achieved an average of 52.44 and 29.77 for chrF++ and BLEU respectively. Our model\footnote{Our code and models are available at \url{https://github.com/maharajbrahma/WAT2024-MultiIndicMT}} is competitive with IndicTransv1 (474M parameter model).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。