arXiv:2410.18697cs.CLcs.AI2024-10NAACL被引 45

评测大模型在文学翻译中的表现,发现人类专家胜过模型。

How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs

  • 构建了包含2000+译文的平行语料库,评估人类与大模型翻译质量。
  • 专业译者评价更准确,简单评分法比复杂标准更有效。
  • 现有大模型翻译更直白单一,仍不及人类译文多样性与美感。

近期研究将文学机器翻译视为新挑战,但其评估仍是开放问题。本文引入LITEVAL-CORPUS,一个段落级双语平行语料库,包含9个机器翻译系统输出及经验证的人类译文,共覆盖4种语言对、超过2000篇译文和13000句评估,耗时4500人时。该语料库使我们能(i)检验不同复杂度下人类评估的一致性与充分性,(ii)比较学生与专业人士的评估差异,(iii)评估基于大模型的自动指标,(iv)直接对比大模型自身表现。结果表明:人类评估的充分性受两个因素影响——评估方案复杂度(越复杂越不充分)与评估者专业水平(越高越准确)。例如,广泛用于非文学翻译的多维质量指标MQM,在学生评估下近60%的人类译文被误判为与机器译文无异或更差;而更简单的最佳最差评分法BWS可识别80%-100%的人类译文。自动指标表现极差,最高准确率仅20%。整体评估显示,已发表的人类译文始终优于大模型译文,即使最新模型也普遍更字面化、缺乏多样性。

原文摘要 · Abstract (English)

Recent research has focused on literary machine translation (MT) as a new challenge in MT. However, the evaluation of literary MT remains an open problem. We contribute to this ongoing discussion by introducing LITEVAL-CORPUS, a paragraph-level parallel corpus containing verified human translations and outputs from 9 MT systems, which totals over 2k translations and 13k evaluated sentences across four language pairs, costing 4.5k C. This corpus enables us to (i) examine the consistency and adequacy of human evaluation schemes with various degrees of complexity, (ii) compare evaluations by students and professionals, assess the effectiveness of (iii) LLM-based metrics and (iv) LLMs themselves. Our findings indicate that the adequacy of human evaluation is controlled by two factors: the complexity of the evaluation scheme (more complex is less adequate) and the expertise of evaluators (higher expertise yields more adequate evaluations). For instance, MQM (Multidimensional Quality Metrics), a complex scheme and the de facto standard for non-literary human MT evaluation, is largely inadequate for literary translation evaluation: with student evaluators, nearly 60% of human translations are misjudged as indistinguishable or inferior to machine translations. In contrast, BWS (BEST-WORST SCALING), a much simpler scheme, identifies human translations at a rate of 80-100%. Automatic metrics fare dramatically worse, with rates of at most 20%. Our overall evaluation indicates that published human translations consistently outperform LLM translations, where even the most recent LLMs tend to produce considerably more literal and less diverse translations compared to humans.

文学翻译大模型评估人类对比质量评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。