arXiv:2502.02577cs.CL2025-02中稿 · MT Summit 2025

对比DeepL与Supertext在长文本翻译中的表现,发现后者更稳定。

A comparison of translation performance between DeepL and Supertext

  • 用全文上下文评估机器翻译质量,避免分段干扰
  • 三组语言方向中Supertext的翻译一致性更优
  • 适合关注真实场景下翻译连贯性的研究者

随着强大机器翻译系统越来越多地基于大语言模型,可靠的品质评估需要能捕捉其利用长上下文的能力。本研究通过评估未分段文本,对比了两个商用MT系统——DeepL与Supertext。在四个语言方向上,由专业译者基于完整文档级上下文评估翻译质量。虽然分段评估显示两者在多数情况下无明显优劣,但文档级分析表明,在四组语言方向中有三组偏好Supertext,说明其在长文本中具备更优的翻译一致性。研究呼吁采用更注重上下文的评估方法,以确保评估结果反映实际使用体验。所有评估数据与脚本已公开于https://github.com/supertext/evaluation_deepl_supertext,供进一步分析与复现。

原文摘要 · Abstract (English)

As strong machine translation (MT) systems are increasingly based on large language models (LLMs), reliable quality benchmarking requires methods that capture their ability to leverage extended context. This study compares two commercial MT systems -- DeepL and Supertext -- by assessing their performance on unsegmented texts. We evaluate translation quality across four language directions with professional translators assessing segments with full document-level context. While segment-level assessments indicate no strong preference between the systems in most cases, document-level analysis reveals a preference for Supertext in three out of four language directions, suggesting superior consistency across longer texts. We advocate for more context-sensitive evaluation methodologies to ensure that MT quality assessments reflect real-world usability. We release all evaluation data and scripts for further analysis and reproduction at https://github.com/supertext/evaluation_deepl_supertext.

机器翻译长文本评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。