arXiv:2510.08870cs.CL2025-10

用质量评估重排序提升文档级机器翻译效果,效果显著且开销低。

Quality Estimation Reranking for Document-Level Translation

  • 基于学习的和大模型的质检指标对生成译文进行重排序
  • 32个候选译文中,最高提升达+5.09(BLEURT-20)
  • 适用于长文档,硬件支持下几乎无额外延迟

质量估计重排序是一种质量感知解码方法,旨在通过评分并选择最佳候选译文来提升机器翻译性能。尽管在句级翻译中已被证明有效,但在日益重要的文档级翻译领域仍研究不足。本文评估了在文档级翻译上使用多种学习型与大语言模型(LLM)驱动的质量估计(QE)指标的重排序表现。结果显示,使用最优学习型指标SLIDE,在仅有两个候选译文时BLEURT-20得分提升+2.00,32个候选时提升+5.09,覆盖仅解码器类LLM与编码器-解码器神经机器翻译(NMT)模型。使用最优的LLM-based指标GEMBA-DA,对应提升为+1.63与+4.30。尽管长输入下增益减弱,但最长文档(512–1024源词元)仍实现+2.34(SLIDE)和+1.40(GEMBA-DA)的提升。结果表明,文档级质量估计具有实际价值,且在合适模型与硬件下运行开销极小。

原文摘要 · Abstract (English)

Quality estimation (QE) reranking is a form of quality-aware decoding which aims to improve machine translation (MT) by scoring and selecting the best candidate from a pool of generated translations. While known to be effective at the sentence level, its application to the increasingly prominent domain of document-level translation remains underexplored. In this work, we evaluate QE reranking performance on document-level (rather than the typical sentence-level) translation, using various learned and large language model (LLM)-based QE metrics. We find that with our best learned metric, SLIDE, BLEURT-20 scores improve by +2.00 with only two candidates, and by +5.09 with 32, across both decoder-only LLM models and encoder-decoder neural machine translation (NMT) models. Using the best LLM-based metric, GEMBA-DA, gains of +1.63 and +4.30 are achieved under the same conditions. Although gains shrink with longer inputs, reranking with 32 candidates yields improvements of +2.34 (SLIDE) and +1.40 (GEMBA-DA) on our longest documents (512-1024 source tokens). These findings demonstrate the practical value of document-level QE, with minimal runtime overhead given suitable translation models and hardware.

机器翻译质量估计重排序大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。