arXiv:2510.19546cs.CL2025-10被引 10

揭示多语言微调中灾难性遗忘的触发条件,关键在模型与数据规模比。

Conditions for Catastrophic Forgetting in Multilingual Translation

  • 通过控制实验发现,模型与数据规模比是遗忘主因。
  • 模型指令跟随能力比架构更影响多语言知识保留。
  • 跨语言对齐能缓解遗忘并促进未见语言迁移。

在多语言基础模型上微调特定语言常引发灾难性遗忘,导致未参与微调的语言性能下降。尽管此现象广泛记录,但现有研究对遗忘发生条件缺乏统一结论。为此,我们以机器翻译为测试平台,系统性地开展实证研究,识别多语言微调中触发遗忘的关键条件。通过在不同模型架构、数据规模和微调方法下的受控实验,我们发现模型与数据的相对规模是遗忘的主要决定因素。此外,我们证明模型的指令遵循能力比其架构更关键于多语言知识保留。与普遍假设相反,参数高效微调在缓解遗忘方面并无明显优于全量微调的优势。最后,我们展示跨语言对齐不仅能减轻遗忘,还能促进对未见目标语言的正向迁移。

原文摘要 · Abstract (English)

Fine-tuning multilingual foundation models on specific languages often induces catastrophic forgetting, degrading performance on languages unseen in fine-tuning. While this phenomenon is widely-documented, the literature presents fragmented results about when forgetting occurs. To address this ambiguity, we conduct a systematic empirical study using machine translation as a testbed to identify the conditions that trigger catastrophic forgetting in multilingual fine-tuning. Through controlled experiments across different model architectures, data scales, and fine-tuning approaches, we reveal that the relative scale between model and data size is a primary determinant of forgetting. Moreover, we demonstrate that a model's instruction-following ability is more critical for retaining multilingual knowledge than its architecture. Contrary to assumptions, parameter-efficient fine-tuning offers no clear advantage over full fine-tuning in mitigating forgetting. Lastly, we show that cross-lingual alignment can mitigate forgetting while also facilitating positive transfer to unseen target languages.

多语言微调遗忘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。