arXiv:2510.25434cs.CL2025-10

对比多种评估方法,发现文本指标难测手语翻译质量。

A Critical Study of Automatic Evaluation in Sign Language Translation

  • 对比六种文本与大模型评估法,测试三种控制条件下的表现
  • 文本指标易受句式变化干扰,大模型更懂语义但有偏见
  • 提示需引入多模态评估,全面衡量手语翻译质量

自动评估指标对推动手语翻译(SLT)发展至关重要。当前的SLT评估指标如BLEU和ROUGE均为文本基底,尚不清楚它们在多大程度上能可靠反映输出质量。为填补这一空白,我们通过分析六种指标——包括BLEU、chrF、ROUGE及BLEURT——与基于大语言模型(LLM)的评估器如G-Eval和GEMBA零样本直接评估,在三种受控条件下:改写、模型输出中的幻觉、句长变化,考察其一致性与鲁棒性。结果表明,基于词汇重叠的指标存在局限;尽管基于大模型的评估器更能捕捉传统指标遗漏的语义等价性,但对大模型改写的翻译可能产生偏差。此外,所有指标均能检测到幻觉,但BLEU过于敏感,而BLEURT与大模型评估器对细微情况则相对宽松。这表明亟需超越纯文本的多模态评估框架,以实现对SLT输出的更全面评价。

原文摘要 · Abstract (English)

Automatic evaluation metrics are crucial for advancing sign language translation (SLT). Current SLT evaluation metrics, such as BLEU and ROUGE, are only text-based, and it remains unclear to what extent text-based metrics can reliably capture the quality of SLT outputs. To address this gap, we investigate the limitations of text-based SLT evaluation metrics by analyzing six metrics, including BLEU, chrF, and ROUGE, as well as BLEURT on the one hand, and large language model (LLM)-based evaluators such as G-Eval and GEMBA zero-shot direct assessment on the other hand. Specifically, we assess the consistency and robustness of these metrics under three controlled conditions: paraphrasing, hallucinations in model outputs, and variations in sentence length. Our analysis highlights the limitations of lexical overlap metrics and demonstrates that while LLM-based evaluators better capture semantic equivalence often missed by conventional metrics, they can also exhibit bias toward LLM-paraphrased translations. Moreover, although all metrics are able to detect hallucinations, BLEU tends to be overly sensitive, whereas BLEURT and LLM-based evaluators are comparatively lenient toward subtle cases. This motivates the need for multimodal evaluation frameworks that extend beyond text-based metrics to enable a more holistic assessment of SLT outputs.

手语翻译自动评估大模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。