arXiv:2409.14335cs.CL2024-09被引 25

用自动校对过滤无效错误,让大模型翻译评估更贴近人工标注。

MQM-APE: Toward High-Quality Error Annotation Predictors with Automatic Post-Editing in LLM Translation Evaluators

  • 通过自动校对原译文中的每处错误,保留影响质量提升的错误。
  • 在8个大模型上均优于现有方法,跨高低资源语言表现稳定。
  • 无需训练,适配各类翻译评估任务,提升反馈可解释性。

大语言模型(LLMs)在机器翻译质量评估中展现出巨大潜力,能提供评分和细粒度反馈。尽管类似GEMBA-MQM的方法在无参考评估中达到顶尖性能,但其预测错误与人工标注不一致,限制了反馈的可解释性。为提升LLM评估器预测错误的质量,我们提出一种通用且无需训练的框架MQM-APE,基于自动后编辑(APE)思想:对每个错误进行自动修正,仅保留对质量改进有贡献的错误。具体地,我们设计LLM扮演三重角色:1)评估者,生成错误标注;2)后编辑者,判断错误是否影响质量提升;3)成对质量验证者,作为错误过滤器。实验表明,该方法在8个大模型上均显著提升错误片段的可靠性和质量,涵盖高/低资源语言。与训练型方法正交,可与特定任务评估器(如Tower)互补,展现广泛适用性。深入分析验证了各模块有效性,为评估器设计与模型选择提供重要洞见。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown significant potential as judges for Machine Translation (MT) quality assessment, providing both scores and fine-grained feedback. Although approaches such as GEMBA-MQM have shown state-of-the-art performance on reference-free evaluation, the predicted errors do not align well with those annotated by human, limiting their interpretability as feedback signals. To enhance the quality of error annotations predicted by LLM evaluators, we introduce a universal and training-free framework, $\textbf{MQM-APE}$, based on the idea of filtering out non-impactful errors by Automatically Post-Editing (APE) the original translation based on each error, leaving only those errors that contribute to quality improvement. Specifically, we prompt the LLM to act as 1) $\textit{evaluator}$ to provide error annotations, 2) $\textit{post-editor}$ to determine whether errors impact quality improvement and 3) $\textit{pairwise quality verifier}$ as the error filter. Experiments show that our approach consistently improves both the reliability and quality of error spans against GEMBA-MQM, across eight LLMs in both high- and low-resource languages. Orthogonal to trained approaches, MQM-APE complements translation-specific evaluators such as Tower, highlighting its broad applicability. Further analysis confirms the effectiveness of each module and offers valuable insights into evaluator design and LLMs selection.

机器翻译大模型评估错误标注自动校对

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。