arXiv:2601.03135cs.CL2026-01被引 2

用合成数据提升美洲原住民语言翻译质量,效果显著。

Improving Indigenous Language Machine Translation with Synthetic Data and Language-Specific Preprocessing

  • 用多语言模型生成合成语料,扩充原住民语言训练集。
  • Guarani-Spanish和Quechua-Spanish翻译任务中chrF++提升明显。
  • 针对黏着语特点优化预处理,适合原住民语言研究者。

低资源原住民语言常缺乏神经机器翻译所需的平行语料。本文提出通过高容量多语言翻译模型生成合成句对,来弥补数据稀缺问题。我们对美洲原住民语言的精选平行语料进行合成数据增强,并在仅使用精选数据与合成增强数据上微调多语言mBART模型。采用chrF++作为主要评估指标,该指标被用于近期美洲NLP共享任务中对黏着语的评测。此外,引入语言特定的预处理方法,包括拼写规范化和抗噪声过滤,以减少语料中的异常。在Guarani-Spanish与Quechua-Spanish翻译任务中,合成数据均带来一致的chrF++性能提升;而对Aymara语言的诊断实验显示,通用预处理方法在高度黏着的语言上存在局限性。

原文摘要 · Abstract (English)

Low-resource indigenous languages often lack the parallel corpora required for effective neural machine translation (NMT). Synthetic data generation offers a practical strategy for mitigating this limitation in data-scarce settings. In this work, we augment curated parallel datasets for indigenous languages of the Americas with synthetic sentence pairs generated using a high-capacity multilingual translation model. We fine-tune a multilingual mBART model on curated-only and synthetically augmented data and evaluate translation quality using chrF++, the primary metric used in recent AmericasNLP shared tasks for agglutinative languages. We further apply language-specific preprocessing, including orthographic normalization and noise-aware filtering, to reduce corpus artifacts. Experiments on Guarani-Spanish and Quechua-Spanish translation show consistent chrF++ improvements from synthetic data augmentation, while diagnostic experiments on Aymara highlight the limitations of generic preprocessing for highly agglutinative languages.

机器翻译合成数据原住民语言多语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。