arXiv:2602.17425cs.CL2026-02被引 3

对比两种评估方法在极低资源翻译中的表现,发现它们各有优势。

Evaluating Extremely Low-Resource Machine Translation: A Comparative Study of ChrF++ and BLEU Metrics

  • 比较了基于词元的BLEU与基于字符的ChrF++在极低资源语言上的表现
  • BLEU虽得分低但能更好反映词汇精度,提升结果可解释性
  • 适合关注低资源翻译评估的NLP研究者和模型开发者

在极低资源语言(ELRL)场景下评估机器翻译质量面临独特挑战,因常用指标如BLEU在高资源环境下有效,但在数据稀缺时常失真。本文对比分析了基于n-gram的BLEU与基于字符的ChrF++在三种极低资源语言(Magahi、Bhojpuri、Chhattisgarhi)中的表现,重点考察大语言模型(LLMs)和神经机器翻译(NMT)系统的输出。研究关注翻译伪影,包括幻觉、重复、源文复制及变音符号(matra)变化。尽管近期研究多仅依赖ChrF++,但结果表明,尽管BLEU绝对分数较低,其仍能提供互补的词汇精确度信息,显著提升评估结果的可解释性。

原文摘要 · Abstract (English)

Evaluating machine translation (MT) quality in extremely low-resource language (ELRL) scenarios poses unique challenges, as widely used metrics such as BLEU, effective in high-resource settings, often misrepresent quality in data-scarce contexts. This work presents a comparative analysis of BLEU, an n-gram-based metric, and ChrF++, a character-based metric, for MT evaluation in ELRL settings. We examine how each metric responds to translation artifacts, including hallucinations, repetition, source-text copying, and diacritic (\textit{matra}) variations across three ELRLs: Magahi, Bhojpuri, and Chhattisgarhi, with a focus on outputs from large language models (LLMs) and neural MT (NMT) systems. While recent work often relies solely on ChrF++, our findings show that BLEU, despite its lower absolute scores, provides complementary lexical-precision insights that improve interpretability.

机器翻译评估指标低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。