arXiv:2412.13881cs.CLcs.AI2024-12

用少量数据提升低资源多语种翻译,探索知识迁移的边界与失效点。

Understanding and Analyzing Model Robustness and Knowledge-Transfer in Multilingual Neural Machine Translation using TX-Ray

  • 以英语为枢纽,通过分步迁移学习实现跨语言知识共享。
  • 在4万句数据上,顺序迁移学习使翻译效果优于基线模型。
  • 剪枝削弱模型鲁棒性,加剧遗忘,不助于泛化能力提升。

神经网络在神经机器翻译(NMT)中已显著超越传统短语翻译方法,但在极低资源条件下的多语种神经机器翻译(MNMT)仍研究不足。本研究探究跨语言知识迁移如何提升此类场景下的翻译表现。基于赫尔辛基自然语言处理组提供的Tatoeba翻译挑战数据集,我们开展英德、英法、英西翻译任务,仅使用极少平行语料建立跨语言映射。不同于依赖特定语对大规模预训练的传统方法,本研究将模型在英语-英语翻译上预训练,设定英语为所有任务的源语言,并采用联合多任务与顺序迁移学习策略进行微调。重点回答三个问题:(1) 跨语言知识迁移如何改善极低资源条件下的MNMT?(2) 剪枝神经元知识对模型泛化性、鲁棒性和灾难性遗忘的影响?(3) TX-Ray如何解释并量化训练模型中的知识迁移?实验采用BLEU-4评分评估,结果显示,在40,000句平行语料上,顺序迁移学习优于基线模型,证明其有效性;但剪枝导致性能下降,加重灾难性遗忘,未能提升鲁棒性或泛化能力。研究揭示了知识迁移在极低资源情境下的潜力与局限。

原文摘要 · Abstract (English)

Neural networks have demonstrated significant advancements in Neural Machine Translation (NMT) compared to conventional phrase-based approaches. However, Multilingual Neural Machine Translation (MNMT) in extremely low-resource settings remains underexplored. This research investigates how knowledge transfer across languages can enhance MNMT in such scenarios. Using the Tatoeba translation challenge dataset from Helsinki NLP, we perform English-German, English-French, and English-Spanish translations, leveraging minimal parallel data to establish cross-lingual mappings. Unlike conventional methods relying on extensive pre-training for specific language pairs, we pre-train our model on English-English translations, setting English as the source language for all tasks. The model is fine-tuned on target language pairs using joint multi-task and sequential transfer learning strategies. Our work addresses three key questions: (1) How can knowledge transfer across languages improve MNMT in extremely low-resource scenarios? (2) How does pruning neuron knowledge affect model generalization, robustness, and catastrophic forgetting? (3) How can TX-Ray interpret and quantify knowledge transfer in trained models? Evaluation using BLEU-4 scores demonstrates that sequential transfer learning outperforms baselines on a 40k parallel sentence corpus, showcasing its efficacy. However, pruning neuron knowledge degrades performance, increases catastrophic forgetting, and fails to improve robustness or generalization. Our findings provide valuable insights into the potential and limitations of knowledge transfer and pruning in MNMT for extremely low-resource settings.

多语种翻译知识迁移低资源模型剪枝

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。