arXiv:2510.00255cs.CLcs.AI2025-10被引 3

用大模型推理能力自动评估翻译质量,更准且可解释。

TASER: Translation Assessment via Systematic Evaluation and Reasoning

  • 用大推理模型分步分析翻译质量,结构化评估
  • 在参考和无参考场景下均达顶尖水平,段级表现最优
  • 结果可解释,适合需透明评估的翻译研究与应用

我们提出TASER(基于系统评估与推理的翻译评估),一种利用大推理模型(LRMs)进行自动化翻译质量评估的指标。TASER借助LRMs的显式推理能力,对翻译质量进行系统性、分步式评估。我们在WMT24指标共享任务中,针对有参考和无参考两种场景评估TASER,表现均达到当前最优。在系统级评估中,TASER在两种设置下均取得最高软配对准确率,优于所有现有指标;在段级评估中,其无参考变体在所有无参考方法中排名第一。实验表明,结构化提示模板比开放式提示更适用于LRMs,而传统LLMs则偏好后者。我们测试了OpenAI的o3模型在不同推理深度下的表现,揭示了推理深度与评估质量的关系。LRM的显式推理过程提升了可解释性,解决了现有自动化评估指标的关键缺陷。结果表明,大推理模型在翻译质量评估中实现显著进步,兼具更高精度与透明评估能力,适用于多种语言对。

原文摘要 · Abstract (English)

We introduce TASER (Translation Assessment via Systematic Evaluation and Reasoning), a metric that uses Large Reasoning Models (LRMs) for automated translation quality assessment. TASER harnesses the explicit reasoning capabilities of LRMs to conduct systematic, step-by-step evaluation of translation quality. We evaluate TASER on the WMT24 Metrics Shared Task across both reference-based and reference-free scenarios, demonstrating state-of-the-art performance. In system-level evaluation, TASER achieves the highest soft pairwise accuracy in both reference-based and reference-free settings, outperforming all existing metrics. At the segment level, TASER maintains competitive performance with our reference-free variant ranking as the top-performing metric among all reference-free approaches. Our experiments reveal that structured prompting templates yield superior results with LRMs compared to the open-ended approaches that proved optimal for traditional LLMs. We evaluate o3, a large reasoning model from OpenAI, with varying reasoning efforts, providing insights into the relationship between reasoning depth and evaluation quality. The explicit reasoning process in LRMs offers interpretability and visibility, addressing a key limitation of existing automated metrics. Our results demonstrate that Large Reasoning Models show a measurable advancement in translation quality assessment, combining improved accuracy with transparent evaluation across diverse language pairs.

机器翻译评估指标大模型推理可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。