arXiv:2410.05183cs.CLcs.AI2024-10EMNLP被引 14

提出可解释的机器翻译评估框架,突破传统相关性评价局限。

Beyond Correlation: Interpretable Evaluation of Machine Translation Metrics

  • 设计双场景评估框架,模拟数据筛选与译文重排序
  • 用准确率、召回率、F1值替代相关性分析,量化评估效果
  • 揭示人工标注数据可靠性问题,警示标准流程缺陷

机器翻译评估指标用于自动判断翻译质量。近年来,这些指标被拓展应用于数据筛选和译文重排序等新场景。然而,多数指标仅输出难以解释的标量分数,限制了设计决策的合理性。以往评估主要依赖与人工判断的相关性,虽有效但无法提供直观性能洞察,尤其在新应用场景下。为此,本文提出一种可解释的评估框架,通过两个代理场景评估指标在数据筛选和重排序中的表现。采用精确率、召回率和F1值衡量指标能力,相较相关性分析更清晰地揭示其优劣。此外,研究发现遵循‘直接评估+标量质量度量’(DA+SQM)指南的人工标注数据,与多维质量度量(MQM)标注存在显著低一致性,质疑其可靠性。

原文摘要 · Abstract (English)

Machine Translation (MT) evaluation metrics assess translation quality automatically. Recently, researchers have employed MT metrics for various new use cases, such as data filtering and translation re-ranking. However, most MT metrics return assessments as scalar scores that are difficult to interpret, posing a challenge to making informed design choices. Moreover, MT metrics' capabilities have historically been evaluated using correlation with human judgment, which, despite its efficacy, falls short of providing intuitive insights into metric performance, especially in terms of new metric use cases. To address these issues, we introduce an interpretable evaluation framework for MT metrics. Within this framework, we evaluate metrics in two scenarios that serve as proxies for the data filtering and translation re-ranking use cases. Furthermore, by measuring the performance of MT metrics using Precision, Recall, and F-score, we offer clearer insights into their capabilities than correlation with human judgments. Finally, we raise concerns regarding the reliability of manually curated data following the Direct Assessments+Scalar Quality Metrics (DA+SQM) guidelines, reporting a notably low agreement with Multidimensional Quality Metrics (MQM) annotations.

机器翻译评估框架可解释性质量度量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。