对比两种评估方法在极低资源翻译中的表现,发现它们各有优势。
Evaluating Extremely Low-Resource Machine Translation: A Comparative Study of ChrF++ and BLEU Metrics
- 比较了基于词元的BLEU与基于字符的ChrF++在极低资源语言上的表现
- BLEU虽得分低但能更好反映词汇精度,提升结果可解释性
- 适合关注低资源翻译评估的NLP研究者和模型开发者
在极低资源语言(ELRL)场景下评估机器翻译质量面临独特挑战,因常用指标如BLEU在高资源环境下有效,但在数据稀缺时常失真。本文对比分析了基于n-gram的BLEU与基于字符的ChrF++在三种极低资源语言(Magahi、Bhojpuri、Chhattisgarhi)中的表现,重点考察大语言模型(LLMs)和神经机器翻译(NMT)系统的输出。研究关注翻译伪影,包括幻觉、重复、源文复制及变音符号(matra)变化。尽管近期研究多仅依赖ChrF++,但结果表明,尽管BLEU绝对分数较低,其仍能提供互补的词汇精确度信息,显著提升评估结果的可解释性。
原文摘要 · Abstract (English)
Evaluating machine translation (MT) quality in extremely low-resource language (ELRL) scenarios poses unique challenges, as widely used metrics such as BLEU, effective in high-resource settings, often misrepresent quality in data-scarce contexts. This work presents a comparative analysis of BLEU, an n-gram-based metric, and ChrF++, a character-based metric, for MT evaluation in ELRL settings. We examine how each metric responds to translation artifacts, including hallucinations, repetition, source-text copying, and diacritic (\textit{matra}) variations across three ELRLs: Magahi, Bhojpuri, and Chhattisgarhi, with a focus on outputs from large language models (LLMs) and neural MT (NMT) systems. While recent work often relies solely on ChrF++, our findings show that BLEU, despite its lower absolute scores, provides complementary lexical-precision insights that improve interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。