arXiv:2501.16581cs.CL2025-01ACL被引 2

让翻译模型适应方言,提升低资源方言的翻译效果

DialUp! Modeling the Language Continuum by Adapting Models to Dialects and Dialects to Models

  • 训练时用合成数据让模型学习方言变化规律,增强泛化能力
  • 推理时调整方言文本,使其更符合模型擅长的语言形式
  • 对低性能方言提升显著,适合资源匮乏语言的研究者

世界上多数语言和方言属于低资源类型,主流机器翻译模型缺乏支持。但许多方言与高资源语言(HRL)关系密切,且在语言上呈现规则性差异。为此,我们提出DialUp:一种训练阶段将预训练模型适配方言数据(M→D)的方法,以及推理阶段将方言数据转换为模型擅长形式(D→M)的干预策略。M→D通过合成数据暴露模型于方言变异机制,提升对未知方言的鲁棒性;D→M则针对已知目标方言处理其语言差异。实验显示,在四个语系的多个方言上取得显著性能提升,另两个语系有适度改善。特征与错误分析表明,基础翻译性能较低的语言变体更受益于该方法。

原文摘要 · Abstract (English)

Most of the world's languages and dialects are low-resource, and lack support in mainstream machine translation (MT) models. However, many of them have a closely-related high-resource language (HRL) neighbor, and differ in linguistically regular ways from it. This underscores the importance of model robustness to dialectal variation and cross-lingual generalization to the HRL dialect continuum. We present DialUp, consisting of a training-time technique for adapting a pretrained model to dialectal data (M->D), and an inference-time intervention adapting dialectal data to the model expertise (D->M). M->D induces model robustness to potentially unseen and unknown dialects by exposure to synthetic data exemplifying linguistic mechanisms of dialectal variation, whereas D->M treats dialectal divergence for known target dialects. These methods show considerable performance gains for several dialects from four language families, and modest gains for two other language families. We also conduct feature and error analyses, which show that language varieties with low baseline MT performance are more likely to benefit from these approaches.

机器翻译方言低资源模型适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。