构建首个非洲语言文档级翻译数据集,助力低资源语言机器翻译
AFRIDOC-MT: Document-level MT Corpus for African Languages
- 收集334篇健康与271篇科技新闻的英-非多语言文档级平行语料
- NLLB-200在标准NMT模型中表现最佳,GPT-4o超越通用大模型
- 发现大模型在非洲语言上存在重复、漏译等现象,适合语言研究者参考
本文提出AFRIDOC-MT,一个覆盖英语及五种非洲语言(阿姆哈拉语、豪萨语、斯瓦希里语、约鲁巴语、祖鲁语)的文档级多语言翻译语料库。数据集包含334篇健康类与271篇信息技术类新闻文档,均由人工从英文翻译至目标语言。通过评估神经机器翻译(NMT)模型和大语言模型(LLM)在句子与伪文档层面的翻译表现,我们对输出进行重新对齐以形成完整文档进行评测。结果表明,标准NMT模型中NLLB-200平均表现最佳,而GPT-4o优于通用大语言模型。微调可显著提升性能,但仅在句子级别训练的模型难以有效泛化到长文档。此外分析发现,部分大模型在非洲语言上存在欠生成、重复词句及偏移翻译等问题。
原文摘要 · Abstract (English)
This paper introduces AFRIDOC-MT, a document-level multi-parallel translation dataset covering English and five African languages: Amharic, Hausa, Swahili, Yorùbá, and Zulu. The dataset comprises 334 health and 271 information technology news documents, all human-translated from English to these languages. We conduct document-level translation benchmark experiments by evaluating neural machine translation (NMT) models and large language models (LLMs) for translations between English and these languages, at both the sentence and pseudo-document levels. These outputs are realigned to form complete documents for evaluation. Our results indicate that NLLB-200 achieved the best average performance among the standard NMT models, while GPT-4o outperformed general-purpose LLMs. Fine-tuning selected models led to substantial performance gains, but models trained on sentences struggled to generalize effectively to longer documents. Furthermore, our analysis reveals that some LLMs exhibit issues such as under-generation, repetition of words or phrases, and off-target translations, especially for African languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。