arXiv:2507.12724cs.CL2025-07被引 1

用推理评估翻译质量,自动打分并排序,效果媲美顶尖模型。

TransEvalnia: Reasoning-based Evaluation and Ranking of Translations

  • 基于多维质量指标,通过大模型推理进行细粒度评分
  • 在英日及多个WMT语言对上表现优于或等同于SOTA模型
  • 输出可解释的评估过程,适合需要透明性的人工智能应用

我们提出TransEvalnia,一种基于提示的翻译评估与排序系统,通过推理完成评估与排序。该系统基于多维质量度量(Multidimensional Quality Metrics)的子集,提供细粒度评估,判断最优翻译,并给出各维度及整体的数值评分。在自建的英日语料以及多个WMT共享任务的语言对上,TransEvalnia的表现与或优于当前最先进的MT-Ranker(Moosa et al. 2024)。使用Anthropic的Claude-3.5-Sonnet和Qwen-2.5-72B-Instruct作为评估大模型时,其输出被人类评审者高度认可,且Sonnet及其他大模型所给分数与人工评分具有强相关性。同时发现本系统及MT-Ranker对翻译顺序敏感,提出缓解位置偏差的方法。所有数据、评估结果、人类标注及代码均已公开。

原文摘要 · Abstract (English)

We present TransEvalnia, a prompting-based translation evaluation and ranking system that uses reasoning in performing its evaluations and ranking. This system presents fine-grained evaluations based on a subset of the Multidimensional Quality Metrics (https://themqm.org/), returns an assessment of which translation it deems the best, and provides numerical scores for the various dimensions and for the overall translation. We show that TransEvalnia performs as well as or better than the state-of-the-art MT-Ranker (Moosa et al. 2024) on our own English-Japanese data as well as several language pairs from various WMT shared tasks. Using Anthropic's Claude-3.5-Sonnet and Qwen-2.5-72B-Instruct as the evaluation LLMs, we show that the evaluations returned are deemed highly acceptable to human raters, and that the scores assigned to the translations by Sonnet, as well as other LLMs, correlate well with scores assigned by the human raters. We also note the sensitivity of our system -- as well as MT-Ranker -- to the order in which the translations are presented, and we propose methods to address this position bias. All data, including the system's evaluation and reasoning, human assessments, as well as code is released.

机器翻译评估系统大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。