arXiv:2601.05864cs.CL2026-01

自动评估指标无法捕捉翻译语境,难真实衡量口译质量。

What do the metrics mean? A critical analysis of the use of Automated Evaluation Metrics in Interpreting

  • 分析多种自动化评估方法的局限性
  • 指出现有指标忽略交际语境,结果不可靠
  • 适合关注口译质量评估标准的研究者

随着远程口译、计算机辅助口译、自动语音翻译及口译虚拟人等技术的发展,快速高效评估口译质量的需求日益增长。为此,研究者提出了多种自动化评估方法。本文系统审视了这些新提出的质量测量手段,探讨其在真实口译场景(无论由人类或机器完成)中的适用性,结论是:当前提出的自动评估指标无法有效考量交际语境,因此单独使用时无法可靠衡量任何口译服务的质量。在口译研究中,口译发生的语境始终是最终分析的核心要素。

原文摘要 · Abstract (English)

With the growth of interpreting technologies, from remote interpreting and Computer-Aided Interpreting to automated speech translation and interpreting avatars, there is now a high demand for ways to quickly and efficiently measure the quality of any interpreting delivered. A range of approaches to fulfil the need for quick and efficient quality measurement have been proposed, each involving some measure of automation. This article examines these recently-proposed quality measurement methods and will discuss their suitability for measuring the quality of authentic interpreting practice, whether delivered by humans or machines, concluding that automatic metrics as currently proposed cannot take into account the communicative context and thus are not viable measures of the quality of any interpreting provision when used on their own. Across all attempts to measure or even categorise quality in Interpreting Studies, the contexts in which interpreting takes place have become fundamental to the final analysis.

口译评估自动度量语境重要

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。