arXiv:2503.22582cs.CL2025-03被引 4

针对低资源语言翻译难题,提出多阶段适配方法提升模型性能。

Beyond Vanilla Fine-Tuning: Leveraging Multistage, Multilingual, and Domain-Specific Methods for Low-Resource Machine Translation

  • 通过持续预训练和中间任务迁移学习,分步增强模型对低资源语言的适应能力。
  • 在少于10万句的数据集上,平均提升1.47 BLEU得分,显著优于单阶段微调。
  • 适合构建低资源语言、多领域场景下的实用机器翻译系统。

微调多语言序列到序列大语言模型(msLLMs)在开发低资源语言(LRLs)神经机器翻译(NMT)系统方面展现出潜力。然而,在训练数据极有限的极端低资源设置下,传统单阶段微调方法表现不佳。本文提出两种适应此类挑战的改进方法:(1) 持续预训练(CPT),即利用领域特定的单语数据进一步训练msLLM,以弥补低资源语言的代表性不足;(2) 中间任务迁移学习(ITTL),通过结合领域内与领域外的平行语料微调模型,提升其跨领域和任务的翻译能力。作为工程应用,这些方法被应用于僧伽罗语、泰米尔语与英语(六种语言对)在特定领域的极低资源设置中(数据集样本数少于10万)。实验表明,相比标准单阶段微调基线,这些方法在所有翻译方向上平均提升1.47 BLEU得分。此外,多模型集成进一步带来额外的性能增益。

原文摘要 · Abstract (English)

Fine-tuning multilingual sequence-to-sequence large language models (msLLMs) has shown promise in developing neural machine translation (NMT) systems for low-resource languages (LRLs). However, conventional single-stage fine-tuning methods struggle in extremely low-resource NMT settings, where training data is very limited. This paper contributes to artificial intelligence by proposing two approaches for adapting msLLMs in these challenging scenarios: (1) continual pre-training (CPT), where the msLLM is further trained with domain-specific monolingual data to compensate for the under-representation of LRLs, and (2) intermediate task transfer learning (ITTL), a method that fine-tunes the msLLM with both in-domain and out-of-domain parallel data to enhance its translation capabilities across various domains and tasks. As an application in engineering, these methods are implemented in NMT systems for Sinhala, Tamil, and English (six language pairs) in domain-specific, extremely low-resource settings (datasets containing fewer than 100,000 samples). Our experiments reveal that these approaches enhance translation performance by an average of +1.47 bilingual evaluation understudy (BLEU) score compared to the standard single-stage fine-tuning baseline across all translation directions. Additionally, a multi-model ensemble further improves performance by an additional BLEU score.

机器翻译低资源多语言微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。