arXiv:2602.04442cs.CLcs.AI2026-02中稿 · EACL 2026被引 1

用合成数据微调模型,提升五种突厥语翻译性能。

No One-Size-Fits-All: Building Systems For Translation to Bashkir, Kazakh, Kyrgyz, Tatar and Chuvash Using Synthetic And Original Data

  • 用合成数据+LoRA微调nllb模型,提升低资源语言翻译效果。
  • 哈萨克语翻译达chrF++ 49.71,巴什基尔语达46.94,表现优异。
  • 适合低资源语言研究者,开源数据与模型权重可用。

我们研究了五组突厥语机器翻译任务:俄语-巴什基尔语、俄语-哈萨克语、俄语-吉尔吉斯语、英语-鞑靼语和英语-楚瓦什语。使用合成数据对nllb-200-distilled-600M模型进行LoRA微调,在哈萨克语上达到chrF++ 49.71,巴什基尔语达46.94。通过检索相似样本提示DeepSeek-V3.2,楚瓦什语实现chrF++ 39.47。鞑靼语的零样本或基于检索的方法获得chrF++ 41.6,吉尔吉斯语零样本方法达45.6。论文发布了数据集和训练权重。

原文摘要 · Abstract (English)

We explore machine translation for five Turkic language pairs: Russian-Bashkir, Russian-Kazakh, Russian-Kyrgyz, English-Tatar, English-Chuvash. Fine-tuning nllb-200-distilled-600M with LoRA on synthetic data achieved chrF++ 49.71 for Kazakh and 46.94 for Bashkir. Prompting DeepSeek-V3.2 with retrieved similar examples achieved chrF++ 39.47 for Chuvash. For Tatar, zero-shot or retrieval-based approaches achieved chrF++ 41.6, while for Kyrgyz the zero-shot approach reached 45.6. We release the dataset and the obtained weights.

机器翻译低资源语言合成数据突厥语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。