arXiv:2603.10775cs.CL2026-03

用大模型生成翻译质量标注,低成本训练出高效质检模型。

Large Language Models as Annotators for Machine Translation Quality Estimation

  • 用大模型生成简化版MQM标注,指导小模型学习。
  • 标注与人工标注相关性高,训练的COMET模型性能优秀。
  • 适合想低成本构建翻译质检系统的研究者和开发者。

大型语言模型(LLMs)在机器翻译质量评估(MTQE)上表现优异,但推理成本过高,难以直接应用。本文提出利用LLMs生成符合MQM风格的标注数据,用于训练COMET模型:遵循Fernandes等人(2023)的研究,我们认为段级标注为大模型提供了充分理由,是实现高质量段级质量评估的关键。我们设计了一种简化的MQM方案,主要限定在顶层类别,以引导大模型选择。提出了基于GPT-4o的系统化提示工程方法,称为PPbMQM(Prompt-Pattern-based-MQM)。实验表明,生成的标注与人工标注高度相关,基于这些数据训练的COMET模型在中英、英德双语段级质检任务上表现具有竞争力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated excellent performance on Machine Translation Quality Estimation (MTQE), yet their high inference costs make them impractical for direct application. In this work, we propose applying LLMs to generate MQM-style annotations for training a COMET model: following Fernandes et al. (2023), we reckon that segment-level annotations provide a strong rationale for LLMs and are key to good segment-level QE. We propose a simplified MQM scheme, mostly restricted to top-level categories, to guide LLM selection. We present a systematic approach for the development of a GPT-4o-based prompt, called PPbMQM (Prompt-Pattern-based-MQM). We show that the resulting annotations correlate well with human annotations and that training COMET on them leads to competitive performance on segment-level QE for Chinese-English and English-German.

大模型翻译质检提示工程COMET

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。