修复4种非洲语言的评估数据集,提升NLP评测准确性
Correcting FLORES Evaluation Dataset for Four African Languages
- 由母语者逐条审校修正数据集中的语言错误
- 纠正后各语言数据质量显著提升,统计差异明显
- 适合关注低资源语言NLP评测与公平性的研究者
本文针对豪萨语、北索托语(塞佩迪)、茨托尼亚语和祖鲁语四种非洲语言的FLORES评估数据集(开发集与测试集)进行了修正。原数据集虽在覆盖低资源语言方面具有开创性,但在这些语言中存在多种不一致和错误,可能影响自然语言处理下游任务(尤其是机器翻译)的评估完整性。通过母语者细致评审,识别并实施了多项修正,提升了数据集整体质量和可靠性。每种语言均总结了发现的错误类型,并提供了修正前后数据的统计对比分析。我们认为,这些修正增强了数据的语言准确性和可信度,有助于更有效评估涉及这四种语言的NLP任务。最后建议未来低资源语言翻译工作应全程吸纳母语者参与,以确保语言准确性和文化相关性。
原文摘要 · Abstract (English)
This paper describes the corrections made to the FLORES evaluation (dev and devtest) dataset for four African languages, namely Hausa, Northern Sotho (Sepedi), Xitsonga, and isiZulu. The original dataset, though groundbreaking in its coverage of low-resource languages, exhibited various inconsistencies and inaccuracies in the reviewed languages that could potentially hinder the integrity of the evaluation of downstream tasks in natural language processing (NLP), especially machine translation. Through a meticulous review process by native speakers, several corrections were identified and implemented, improving the overall quality and reliability of the dataset. For each language, we provide a concise summary of the errors encountered and corrected and also present some statistical analysis that measures the difference between the existing and corrected datasets. We believe that our corrections improve the linguistic accuracy and reliability of the data and, thereby, contribute to a more effective evaluation of NLP tasks involving the four African languages. Finally, we recommend that future translation efforts, particularly in low-resource languages, prioritize the active involvement of native speakers at every stage of the process to ensure linguistic accuracy and cultural relevance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。