构建卡拉卡尔帕克语翻译数据集与模型,助力低资源语言技术发展
Open Language Data Initiative: Advancing Low-Resource Machine Translation for Karakalpak
- 构建10万句对的多语种平行语料库,覆盖乌兹别克、俄、英与卡拉卡尔帕克语
- 基于新数据训练模型,在基准测试中超越已有方法
- 开源数据与模型,推动低资源语言NLP研究
本研究为卡拉卡尔帕克语贡献了多项成果:一个已翻译至卡拉卡尔帕克语的FLORES+开发测试集,以及各含10万句对的乌兹别克-卡拉卡尔帕克、俄-卡拉卡尔帕克和英-卡拉卡尔帕克平行语料库,并开源了这些语言间的微调神经翻译模型。实验对比了不同模型变体与训练方法,证明其在现有基线上的性能提升。该工作作为开放语言数据倡议(OLDI)共享任务的一部分,旨在提升卡拉卡尔帕克语的机器翻译能力,推动自然语言处理技术中的语言多样性。
原文摘要 · Abstract (English)
This study presents several contributions for the Karakalpak language: a FLORES+ devtest dataset translated to Karakalpak, parallel corpora for Uzbek-Karakalpak, Russian-Karakalpak and English-Karakalpak of 100,000 pairs each and open-sourced fine-tuned neural models for translation across these languages. Our experiments compare different model variants and training approaches, demonstrating improvements over existing baselines. This work, conducted as part of the Open Language Data Initiative (OLDI) shared task, aims to advance machine translation capabilities for Karakalpak and contribute to expanding linguistic diversity in NLP technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。