arXiv:2412.03152cs.CLcs.AI2024-12ACL被引 3

提出评估机器翻译自动指标公平性的新方法

A Measure of the System Dependence of Automated Metrics

  • 设计新度量方法评估指标对不同系统的一致性
  • 揭示现有指标在不同系统间存在不公平表现
  • 适合关注评估可靠性的NLP研究者使用

机器翻译的自动评估指标取得了显著进展,目标是替代昂贵且耗时的人工评估。这些指标通常通过与人类判断的相关性来评估,以捕捉人类评分与指标评分之间的单调关系。然而,我们主张确保指标对所有系统公平且一致同样重要。本文提出一种方法来评估这一方面,揭示现有指标在不同系统间的公平性差异。

原文摘要 · Abstract (English)

Automated metrics for Machine Translation have made significant progress, with the goal of replacing expensive and time-consuming human evaluations. These metrics are typically assessed by their correlation with human judgments, which captures the monotonic relationship between human and metric scores. However, we argue that it is equally important to ensure that metrics treat all systems fairly and consistently. In this paper, we introduce a method to evaluate this aspect.

机器翻译评估指标公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。