用大模型微调提升低资源印地语族语言翻译效果
Yes-MT's Submission to the Low-Resource Indic Language Translation Shared Task in WMT 2024
- 对比多种预训练模型在多/单语环境下的微调策略
- 基于Llama 3的LoRA微调在各语言上表现最佳
- 适合关注低资源语言翻译的NLP研究者
本文介绍Yes-MT团队在WMT 2024低资源印地语族语言翻译共享任务中的系统,聚焦英语与阿萨姆语、米佐语、卡西语及曼尼普里语之间的翻译。实验对比了mT5、IndicBart等多语言模型的微调,LoRA微调IndicTrans2,以及Llama 3、Mixtral 8x7b等大模型的零样本与少样本提示方法。还尝试了从头训练Transformer模型。评估基于WMT23测试集,使用SacreBLEU和CHRF指标,结果表明低资源翻译挑战显著,而大模型微调(尤其是Llama 3)展现出巨大潜力。
原文摘要 · Abstract (English)
This paper presents the systems submitted by the Yes-MT team for the Low-Resource Indic Language Translation Shared Task at WMT 2024 (Pakray et al., 2024), focusing on translating between English and the Assamese, Mizo, Khasi, and Manipuri languages. The experiments explored various approaches, including fine-tuning pre-trained models like mT5 (Xue et al., 2020) and IndicBart (Dabre et al., 2021) in both multilingual and monolingual settings, LoRA (Hu et al., 2021) fine-tuning IndicTrans2 (Gala et al., 2023), zero-shot and few-shot prompting (Brown, 2020) with large language models (LLMs) like Llama 3 (Dubey et al., 2024) and Mixtral 8x7b (Jiang et al., 2024), LoRA supervised fine-tuning of Llama 3 (Mecklenburg et al., 2024), and training Transformer models (Vaswani, 2017) from scratch. The results were evaluated on the WMT23 Low-Resource Indic Language Translation Shared Task test data using SacreBLEU (Post, 2018) and CHRF (Popovic, 2015), highlighting the challenges of low-resource translation and the potential of LLMs for these tasks, particularly with fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。