arXiv:2609.03734cs.CLcs.AI2026-09

提出新评估标准,更真实反映手语翻译能力

Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks

论文配图:Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks
图 1 · 摘自论文原文
  • 用开放权重大模型问答测试替代传统BLEU-4
  • 发现现有模型在关键内容保留上差距明显
  • 适合关注手语理解真实水平的研究者

BLEU-4是手语翻译(SLT)的标准评估指标,但其基于口语的评价方式难以准确衡量手语能力。在多模态、低资源的SLT场景中,模型可能利用虚假相关性和口语先验,而非学习真正的手语表示。本文在Phoenix-2014T和CSL-Daily数据集上评估了六种SLT模型的时空理解与BLEU-4的关系,发现BLEU-4提升并不意味着手语理解更好。为此,本文引入受语言学习评估启发的新型评估方法:使用开放权重大模型进行问答测试,以衡量关键内容保留。该方法与人工评分更一致,且比BLEU-4高出六至七倍的抗改写能力。应用于SLT后,该协议揭示了不同系统间的真实差异:五个无词元(gloss-free)系统在Phoenix-2014T上表现相近,而有词元监督的系统高出9.3分,这一差距在BLEU-4中无法体现。

原文摘要 · Abstract (English)

BLEU-4 is the standard metric for evaluating sign language translation (SLT), but spoken-language metrics may not adequately reflect sign language proficiency. The multimodal, low-resource context of SLT allows models to exploit spurious correlations and spoken-language priors, rather than learning stronger sign representations. In this paper, we evaluate the relationship between spatio-temporal understanding and BLEU-4 across six SLT models on Phoenix-2014T and CSL-Daily, showing that gains in BLEU-4 are not on their own evidence of better sign language understanding. This work introduces an alternative inspired by language-learning assessment, using an open-weight-LLM QA protocol that measures salient content preservation. It aligns more closely with human rankings and is six to seven times more paraphrase-invariant than BLEU-4. Applied to SLT, this protocol targets content transfer, is more robust to train-test overlap, and gives a different picture of the field: the five gloss-free systems are largely within noise of one another on Phoenix-2014T, while the gloss-supervised system stands 9.3 points higher, a gap invisible to BLEU-4.

手语翻译评估基准大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。