arXiv:2603.14873cs.CL2026-03中稿 · AfricaNLP 2026被引 1

为濒危语言埃菲克语构建翻译系统,助力文化数字化包容

Developing an English-Efik Corpus and Machine Translation System for Digitization Inclusion

  • 用1.3万句社区收集的英埃菲克双语语料微调模型
  • NLLB-200模型达到31.21的回译BLEU分,优于mT5
  • 推动低资源语言技术普惠,适合关注文化保护的研究者

低资源语言是人类历史与文化多样性的重要载体,但其在现代自然语言处理系统中仍严重缺失。尽管斯瓦希里语、约鲁巴语等非洲主流语言已有进展,埃菲克语等小型土著语言在机器翻译研究中依然被忽视。本研究基于13,865句对的社区共建平行语料,评估了先进多语言神经机器翻译模型在英埃菲克语翻译中的表现。我们对mT5和NLLB200模型进行了微调。结果显示,NLLB-200表现更优,英文到埃菲克语的BLEU得分为26.64,埃菲克语到英文为31.21,对应chrF分数分别为51.04和47.92,表明流畅性与语义保真度更高。研究证明了为低资源语言开发实用翻译工具的可行性,并强调包容性数据实践与文化相关评估在实现公平自然语言处理中的重要性。

原文摘要 · Abstract (English)

Low-resource languages serve as invaluable repositories of human history, preserving cultural and intellectual diversity. Despite their significance, they remain largely absent from modern natural language processing systems. While progress has been made for widely spoken African languages such as Swahili, Yoruba, and Amharic, smaller indigenous languages like Efik continue to be underrepresented in machine translation research. This study evaluates the effectiveness of state-of-the-art multilingual neural machine translation models for English-Efik translation, leveraging a small-scale, community-curated parallel corpus of 13,865 sentence pairs. We fine-tuned both the mT5 multilingual model and the NLLB200 model on this dataset. NLLB-200 outperformed mT5, achieving BLEU scores of 26.64 for English-Efik and 31.21 for Efik-English, with corresponding chrF scores of 51.04 and 47.92, indicating improved fluency and semantic fidelity. Our findings demonstrate the feasibility of developing practical machine translation tools for low-resource languages and highlight the importance of inclusive data practices and culturally grounded evaluation in advancing equitable NLP.

机器翻译低资源语言文化保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。