提出更可靠的评估方法,用于衡量机器翻译错误检测工具的性能。
Span-Level Machine Translation Meta-Evaluation
- 设计新评估策略,解决错误定位评估中的歧义问题
- 发现常用方法会导致结果偏差,影响评估可信度
- 适合关注翻译质量评估的科研人员与开发者
近年来,机器翻译(MT)及自动评估技术取得显著进展,使众多新应用成为可能。自动评估已从输出单一质量分数发展为精确识别翻译错误,并分配错误类别和严重程度。然而,如何可靠评估具备错误检测能力的自动评估工具仍缺乏成熟方法。本文研究了不同实现方式的跨度级精确率、召回率和F分数,表明看似相似的方法可能产生截然不同的排名结果,且某些广泛使用的技术不适用于评估机器翻译错误检测。为此,我们提出“带部分重叠与部分评分”的匹配策略(MPP),结合微平均法作为稳健的元评估方案,并公开发布代码。最后,利用MPP对当前最先进的机器翻译错误检测技术进行评估。
原文摘要 · Abstract (English)
Machine Translation (MT) and automatic MT evaluation have improved dramatically in recent years, enabling numerous novel applications. Automatic evaluation techniques have evolved from producing scalar quality scores to precisely locating translation errors and assigning them error categories and severity levels. However, it remains unclear how to reliably measure the evaluation capabilities of auto-evaluators that do error detection, as no established technique exists in the literature. This work investigates different implementations of span-level precision, recall, and F-score, showing that seemingly similar approaches can yield substantially different rankings, and that certain widely-used techniques are unsuitable for evaluating MT error detection. We propose "match with partial overlap and partial credit" (MPP) with micro-averaging as a robust meta-evaluation strategy and release code for its use publicly. Finally, we use MPP to assess the state of the art in MT error detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。