针对非洲低资源语言翻译,构建了超大规模评测数据集并提出新评估模型。
SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced African Languages?
- 基于7.3万条人工标注,构建14种非洲语言的机器翻译评测数据集
- 新模型在特维语、卢奥语等低资源语言上超越现有方法,接近顶尖大模型表现
- 开源全部数据与模型,助力非洲语言NLP研究
评估低资源非洲语言的机器翻译质量仍面临重大挑战,现有度量工具常因语言覆盖有限且在低资源环境下表现不佳。尽管近期如AfriCOMET等方法已有所改进,但仍受限于小规模评测集、缺乏面向非洲语言的公开训练数据,以及在极端低资源场景下表现不稳定。本文构建了SSA-MTE,一个涵盖14种非洲语言对、来自新闻领域的大型人工标注机器翻译评估数据集,包含超过7.3万条句级标注,覆盖多样化的MT系统。基于此,我们开发了参考依赖型和参考无依赖型的改进度量模型SSA-COMET与SSA-COMET-QE。同时,我们还评估了GPT-4o、Claude-3.7和Gemini 2.5 Pro等先进大语言模型的提示调用方法。实验表明,SSA-COMET显著优于AfriCOMET,且在特维语、卢奥语和约鲁巴语等低资源语言上表现可与最强的大型语言模型Gemini 2.5 Pro相媲美。所有资源均以开源许可证发布,支持后续研究。
原文摘要 · Abstract (English)
Evaluating machine translation (MT) quality for under-resourced African languages remains a significant challenge, as existing metrics often suffer from limited language coverage and poor performance in low-resource settings. While recent efforts, such as AfriCOMET, have addressed some of the issues, they are still constrained by small evaluation sets, a lack of publicly available training data tailored to African languages, and inconsistent performance in extremely low-resource scenarios. In this work, we introduce SSA-MTE, a large-scale human-annotated MT evaluation (MTE) dataset covering 14 African language pairs from the News domain, with over 73,000 sentence-level annotations from a diverse set of MT systems. Based on this data, we develop SSA-COMET and SSA-COMET-QE, improved reference-based and reference-free evaluation metrics. We also benchmark prompting-based approaches using state-of-the-art LLMs like GPT-4o, Claude-3.7 and Gemini 2.5 Pro. Our experimental results show that SSA-COMET models significantly outperform AfriCOMET and are competitive with the strongest LLM Gemini 2.5 Pro evaluated in our study, particularly on low-resource languages such as Twi, Luo, and Yoruba. All resources are released under open licenses to support future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。