arXiv:2412.17537cs.CL2024-12EMNLP被引 7

探究机器翻译领域适应中遗忘的根源与数据关系。

Domain adapted machine translation: What does catastrophic forgetting forget and why?

  • 分析适应数据对模型遗忘的影响机制。
  • 发现遗忘程度与目标词汇覆盖率相关。
  • 为更优领域适配提供理论依据,适合研究者参考。

神经机器翻译(NMT)模型可通过领域适应进行专业化,通常通过在特定数据集上微调实现。这一过程可能引发灾难性遗忘:通用翻译质量迅速下降。尽管遗忘现象已被广泛观察并提出多种缓解方法,其成因及与适应数据的关系仍不明确。本文首次系统探究了遗忘的内容及其原因,考察了遗忘与域内数据之间的关系,发现遗忘程度与数据的目标词汇覆盖率密切相关。研究结果为更明智的NMT领域适应提供了新思路。

原文摘要 · Abstract (English)

Neural Machine Translation (NMT) models can be specialized by domain adaptation, often involving fine-tuning on a dataset of interest. This process risks catastrophic forgetting: rapid loss of generic translation quality. Forgetting has been widely observed, with many mitigation methods proposed. However, the causes of forgetting and the relationship between forgetting and adaptation data are under-explored. This paper takes a novel approach to understanding catastrophic forgetting during NMT adaptation by investigating the impact of the data. We provide a first investigation of what is forgotten, and why. We examine the relationship between forgetting and the in-domain data, and show that the amount and type of forgetting is linked to that data's target vocabulary coverage. Our findings pave the way toward better informed NMT domain adaptation.

机器翻译领域适应遗忘机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。