arXiv:2412.18190cs.CLcs.AI2024-12中稿 · the 29th Annual Me…

对比多种评估指标在日英聊天翻译中的表现,发现传统方法仍可用,神经方法更贴近人工判断。

An Analysis on Automated Metrics for Evaluating Japanese-English Chat Translation

  • 对比了BLEU、TER、BERTScore和COMET等指标在聊天翻译中的表现
  • 神经类指标与人工评分相关性更高,COMET表现最佳
  • 即使最优指标在日语零代词句翻译中仍存在评估困难

本文分析了传统基准指标(如BLEU、TER)和基于神经网络的方法(如BERTScore、COMET)在评估日英聊天翻译中多个神经机器翻译模型性能时的表现,并与人工标注评分进行对比。结果表明,在排序NMT模型优劣方面,所有指标均保持一致,说明传统指标因快速简便仍具实用性。而在与人工判断的相关性上,神经基指标显著优于传统指标,其中COMET在聊天翻译任务中与人工评分相关性最高。然而,研究也发现,即使表现最好的指标,在评估包含日语回指零代词的英文译文时仍存在明显局限。

原文摘要 · Abstract (English)

This paper analyses how traditional baseline metrics, such as BLEU and TER, and neural-based methods, such as BERTScore and COMET, score several NMT models performance on chat translation and how these metrics perform when compared to human-annotated scores. The results show that for ranking NMT models in chat translations, all metrics seem consistent in deciding which model outperforms the others. This implies that traditional baseline metrics, which are faster and simpler to use, can still be helpful. On the other hand, when it comes to better correlation with human judgment, neural-based metrics outperform traditional metrics, with COMET achieving the highest correlation with the human-annotated score on a chat translation. However, we show that even the best metric struggles when scoring English translations from sentences with anaphoric zero-pronoun in Japanese.

机器翻译评估指标日英翻译COMET

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。