arXiv:2505.18905cs.CL2025-05

首个公开的英语-克佩勒语翻译数据集,助力低资源语言技术发展。

Building a Functional Machine Translation Corpus for Kpelle

  • 构建2000+句对的英克佩勒语平行语料库,覆盖日常、宗教与教育文本。
  • 微调NLLB模型后,克佩勒语转英语方向达30分BLEU,验证数据增强有效性。
  • 推动语音识别与语言建模等任务,适合低资源语言研究者参考。

本文首次发布公开可用的英语-克佩勒语机器翻译数据集,包含超过2000个句子对,源自日常交流、宗教文本和教育材料。通过在该数据集两个版本上微调Meta的No Language Left Behind (NLLB) 模型,克佩勒语到英语方向最高获得30分的BLEU分数,证明了数据增强的有效性。研究结果与NLLB-200在其他非洲语言上的基准表现一致,表明尽管克佩勒语为低资源语言,仍具备竞争性性能潜力。该数据集不仅支持机器翻译,还可用于语音识别与语言建模等更广泛的自然语言处理任务。文章最后提出未来扩展路线图,强调拼写一致性、社区参与验证及跨学科合作,以推动克佩勒语及其他低资源曼德语族语言的技术包容性发展。

原文摘要 · Abstract (English)

In this paper, we introduce the first publicly available English-Kpelle dataset for machine translation, comprising over 2000 sentence pairs drawn from everyday communication, religious texts, and educational materials. By fine-tuning Meta's No Language Left Behind(NLLB) model on two versions of the dataset, we achieved BLEU scores of up to 30 in the Kpelle-to-English direction, demonstrating the benefits of data augmentation. Our findings align with NLLB-200 benchmarks on other African languages, underscoring Kpelle's potential for competitive performance despite its low-resource status. Beyond machine translation, this dataset enables broader NLP tasks, including speech recognition and language modelling. We conclude with a roadmap for future dataset expansion, emphasizing orthographic consistency, community-driven validation, and interdisciplinary collaboration to advance inclusive language technology development for Kpelle and other low-resourced Mande languages.

机器翻译低资源语言数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。