arXiv:2503.20007cs.CLcs.LG2025-03NAACL被引 5

无标注数据下构建混用哈俄语的机器翻译模型

Low-resource Machine Translation for Code-switched Kazakh-Russian Language Pair

  • 通过合成数据生成构建模型,无需真实标注数据
  • 达到16.48 BLEU,接近商用系统性能
  • 首个哈俄混用双语语料库,适合低资源语言研究

低资源语言对的机器翻译极具挑战性,而说话者使用语码转换时难度进一步加大。本文提出一种无需标注数据即可构建哈俄语混用语言对机器翻译模型的方法,基于合成数据生成。此外,我们发布了首个哈俄混用平行语料库,并提供了评估结果:所提模型在测试中取得16.48的BLEU分数,几乎达到现有商用系统水平,并在人工评估中表现更优。

原文摘要 · Abstract (English)

Machine translation for low resource language pairs is a challenging task. This task could become extremely difficult once a speaker uses code switching. We propose a method to build a machine translation model for code-switched Kazakh-Russian language pair with no labeled data. Our method is basing on generation of synthetic data. Additionally, we present the first codeswitching Kazakh-Russian parallel corpus and the evaluation results, which include a model achieving 16.48 BLEU almost reaching an existing commercial system and beating it by human evaluation.

机器翻译低资源语码转换混合语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。