arXiv:2507.12672cs.CL2025-07被引 1

首个开源车臣语-俄语机器翻译系统,助力濒危语言保护

The first open machine translation system for the Chechen language

  • 基于NLLB-200微调,构建车臣语与俄语互译模型
  • 俄语→车臣语翻译BLEU 8.34,车臣语→俄语BLEU 20.89
  • 开源平行语料库与多语言编码器,适合语言保护研究者

我们推出了首个开源的车臣语-俄语机器翻译模型及训练评估数据集。通过在多语言翻译大模型NLLB-200上进行微调,实现了车臣语与俄语之间的互译。模型在俄语到车臣语方向的BLEU分数为8.34,ChrF++为34.69;反向翻译(车臣语到俄语)的BLEU为20.89,ChrF++为44.55。该系统配套发布了平行词汇、短语和句子语料库,以及适配车臣语的多语言句向量编码器。

原文摘要 · Abstract (English)

We introduce the first open-source model for translation between the vulnerable Chechen language and Russian, and the dataset collected to train and evaluate it. We explore fine-tuning capabilities for including a new language into a large language model system for multilingual translation NLLB-200. The BLEU / ChrF++ scores for our model are 8.34 / 34.69 and 20.89 / 44.55 for translation from Russian to Chechen and reverse direction, respectively. The release of the translation models is accompanied by the distribution of parallel words, phrases and sentences corpora and multilingual sentence encoder adapted to the Chechen language.

机器翻译濒危语言开源模型NLLB

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。