提出新评估方法,让机器生成的手语更符合人类判断。
BackTranslation2.0 -- A Linguistically Motivated Metric to Assess Sign Language Production

- 用多工具协同框架评估手语语法、发音、动作流畅性等四维度。
- 在英式手语数据集上与真人评分高度一致,相关性显著提升。
- 适合手语生成系统研发者及聋人语言专家参考使用。
手语是数百万聋人主要交流方式,但现有生成手语的评估指标仍过于简单且与人类判断不符。本文提出BackTranslation2.0,一种基于语言学原理的文本转手语评估方法,超越传统回译。该方法采用确定性流水线架构,集成多个专用工具,从语法正确性、音位准确性、动作流畅性和生成保真度四个维度进行评估。工具输出不独立处理:通过大语言模型(LLM)驱动的交叉比对模块,检查各工具间一致性并对照语言预期,实现对语法、音位和动作层面证据的结构化推理。最终得分由经过验证的工具输出通过确定性加权公式计算得出。为验证该方法,我们引入并评估了一个英式手语(BSL)数据集,该数据集由语言学家与聋人专家协作设计,涵盖上述四项质量维度的人类评分,与六种基线指标进行对比。结果表明,BackTranslation2.0在所有维度上均与人类评分呈现强相关性,提供了一个更全面、可解释且语言学基础坚实的评估框架。
原文摘要 · Abstract (English)
Sign Languages (SLs) are the primary means of communication for millions of deaf individuals, yet existing evaluation metrics for generated SL remain simplistic and poorly aligned with human judgements. We introduce BackTranslation2.0, a linguistically grounded evaluation metric for text-to-sign translation that moves beyond naïve backtranslation. Our approach adopts an agentic framework in which a deterministic pipeline orchestrates a suite of specialised tools to assess four scoring dimensions - grammatical correctness, phonological accuracy, motion fluency, and generation fidelity - aligned with human rater assessments. Tool outputs are not treated independently: a set of large language model (LLM)-based cross-referential comparison modules evaluates consistency across tools and checks outputs against linguistic expectations, enabling structured reasoning over grammatical, phonological, and motion-level evidence. Final dimension scores are computed through deterministic weighted formulas over validated tool outputs. To validate BackTranslation2.0, we introduce and evaluate on a British Sign Language (BSL) dataset rated in a human rater study across the same quality dimensions, following a protocol developed in collaboration between linguists and deaf experts, benchmarking against six baseline metrics. Our method demonstrates strong correlation with human judgements across all dimensions, providing a more comprehensive, interpretable, and linguistically principled evaluation framework for sign language production systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。