arXiv:2510.24664cs.CL2025-10ACL被引 1

通过重标注提升机器翻译人工评估质量,发现并修正初评遗漏错误。

MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation

  • 让标注员重新审阅已有翻译评估结果,修正初评疏漏。
  • 重标注后评估质量显著提升,主要因发现原初评遗漏错误。
  • 适合需要高精度翻译评估的研究者与评测团队使用。

机器翻译的人工评估正面临模型质量提升带来的挑战:随着模型性能提高,现有评估方法需优化以避免质量提升被评估噪声掩盖。为此,我们实验了一种两阶段的当前最优评估范式(MQM)变体,称为MQM重标注。在此设置中,一名标注员会审阅并修改一组已有的MQM标注,这些标注可能来自自己、其他标注员或自动标注系统。实验表明,标注员在重标注过程中的行为符合预期目标,且重标注显著提升了标注质量,主要得益于发现了首次标注时被忽略的错误。

原文摘要 · Abstract (English)

Human evaluation of machine translation is in an arms race with translation model quality: as our models get better, our evaluation methods need to be improved to ensure that quality gains are not lost in evaluation noise. To this end, we experiment with a two-stage version of the current state-of-the-art translation evaluation paradigm (MQM), which we call MQM re-annotation. In this setup, an MQM annotator reviews and edits a set of pre-existing MQM annotations, that may have come from themselves, another human annotator, or an automatic MQM annotation system. We demonstrate that rater behavior in re-annotation aligns with our goals, and that re-annotation results in higher-quality annotations, mostly due to finding errors that were missed during the first pass.

机器翻译评估方法标注优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。