arXiv:2603.04083cs.CL2026-03中稿 · the 2026 Language …

研究大模型时代机器翻译质量预测新方法

Hindsight Quality Prediction Experiments in Multi-Candidate Human-Post-Edited Machine Translation

  • 基于多候选翻译数据集,对比源端难度与候选端质检模型
  • 大模型提升文档级翻译质量,但削弱传统预测可靠性
  • 适合关注MT后编辑与LLM应用的研究者

本文研究机器翻译质量预测的两种互补范式:源端难度预测与候选端质量估计(QE)。随着大型语言模型(LLMs)在翻译流程中的快速应用,其对既有质量预测方法的影响尚未充分探索。我们通过一系列“事后”实验,在一个来自真实机翻后编辑(MTPE)项目的独特多候选数据集上开展研究。该数据集包含超过6,000个英文源片段,每个片段对应来自多种传统神经机器翻译系统和先进LLMs的九个译文候选,均以单一人工后编辑参考为基准进行评估。采用肯德尔等级相关系数,我们评估了源端难度指标、候选端QE模型及位置启发式方法在两个金标准评分上的预测能力:以TER作为后编辑工作量的代理,以COMET作为人类判断的代理。研究发现,向大模型的架构转变改变了现有质量预测方法的可靠性,同时缓解了此前文档级翻译中的若干挑战。

原文摘要 · Abstract (English)

This paper investigates two complementary paradigms for predicting machine translation (MT) quality: source-side difficulty prediction and candidate-side quality estimation (QE). The rapid adoption of Large Language Models (LLMs) into MT workflows is reshaping the research landscape, yet its impact on established quality prediction paradigms remains underexplored. We study this issue through a series of "hindsight" experiments on a unique, multi-candidate dataset resulting from a genuine MT post-editing (MTPE) project. The dataset consists of over 6,000 English source segments with nine translation hypotheses from a diverse set of traditional neural MT systems and advanced LLMs, all evaluated against a single, final human post-edited reference. Using Kendall's rank correlation, we assess the predictive power of source-side difficulty metrics, candidate-side QE models and position heuristics against two gold-standard scores: TER (as a proxy for post-editing effort) and COMET (as a proxy for human judgment). Our findings highlight that the architectural shift towards LLMs alters the reliability of established quality prediction methods while simultaneously mitigating previous challenges in document-level translation.

机器翻译质量预测大模型后编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。