arXiv:2606.22269cs.CLcs.AI2026-06

评测大模型在两种西非低资源语言上的翻译表现,发现指标可靠性差异大。

Evaluating Large Language Models for Hausa and Fongbe Machine Translation: Benchmarks, Failures, and Metric Reliability

  • 对比四种大模型在豪萨语和丰贝语上的翻译质量,按规模逐步测试
  • 豪萨语人类评分4.0-4.5,丰贝语仅1.0-2.2,存在3倍BLEU差距
  • 神经类评估指标易失效,需多指标验证,且至少2500句才稳定

我们评估当前大语言模型在英语到豪萨语及英语到丰贝语翻译中的表现——这两种语言分属非洲-亚细亚与尼日-刚果语系,具有显著语言类型差异。在500至10,000句的渐进规模下,使用标准自动评估指标(BLEU、chrF++、TER、COMET、BERTScore)并结合母语者评分进行验证。结果显示:第一,翻译质量因语言而异,豪萨语达到可接受水平(人类评分4.0–4.5/5),而丰贝语表现差(1.0–2.2/5),所有系统间均存在约3倍的BLEU差距;第二,模型排名随语言变化,Gemini在丰贝语中领先,而GPT-4o在豪萨语中胜出,表明一种语言的表现无法预测另一种;第三,评估指标与人类判断的相关性差异巨大,丰贝语中相关系数rho=1.0,豪萨语中仅rho=0.5,且人类偏好GPT-4o,但所有自动指标均将Claude排在首位。进一步发现,如BERTScore等神经指标在两类语言中均出现嵌入坍塌(同类相似度>0.99),削弱其区分能力。研究建议对低资源非洲语言采用多指标评估,并强调最小样本量n=2,500以获得稳定排名,小样本易产生虚假结论。

原文摘要 · Abstract (English)

We investigate the translation quality of current large language models (LLMs) for English-to-Hausa and English-to-Fongbe - two typologically distinct West African languages from the Afroasiatic and Niger-Congo families respectively - and evaluate whether standard automatic metrics reliably reflect human judgment for these low-resource languages. We evaluate four models (GPT-4o Mini, Claude Sonnet 4, Gemini 2.5 Flash, and Qwen2.5-7B) at progressive scales (500 to 10,000 sentences) using automatic metrics (BLEU, chrF++, TER, COMET, BERTScore) validated against native-speaker judgment. Our results reveal three key findings. First, translation quality varies substantially by language: Hausa achieves acceptable quality (human scores 4.0-4.5/5) while Fongbe achieves poor quality (1.0-2.2/5), with a consistent 3x BLEU gap across all systems. Second, model rankings differ by language - Gemini leads for Fongbe while GPT-4o leads for Hausa by human evaluation - indicating that performance on one low-resource African language does not predict performance on another. Third, metric-human correlation varies dramatically: perfect rank correlation for Fongbe (rho=1.0) but weak correlation for Hausa (rho=0.5), where human evaluators preferred GPT-4o despite all automatic metrics ranking Claude first. We further show that neural metrics like BERTScore exhibit embedding collapse (within-language similarity >0.99) for both languages, limiting their ability to differentiate translation quality. Based on these findings, we recommend multi-metric evaluation for low-resource African languages, with particular caution when interpreting neural metrics. We establish that minimum sample sizes of n=2,500 sentences are required for stable system rankings, as smaller samples produced artifact findings that reversed at scale.

机器翻译低资源语言大模型评测评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。