arXiv:2509.03791cs.CLcs.AI2025-09被引 6

用语义嵌入评估手语生成,更准更可靠。

SiLVERScore: Semantically-Aware Embeddings for Sign Language Generation Evaluation

  • 构建联合嵌入空间,直接评估手语生成质量
  • 在两个数据集上准确区分正确与随机对,AUC达0.99
  • 适合研究手语生成与评估的学者使用

手语生成评估通常采用回译法:先将生成的手语识别为文本,再用文本指标对比参考文本。但这种两步流程存在歧义:既无法捕捉手语的多模态特性(如面部表情、空间语法、语调),也难以判断误差源于生成模型还是翻译系统。本文提出SiLVERScore,一种基于语义感知嵌入的评估方法,在联合嵌入空间中评估手语生成。贡献包括:(1)揭示现有评估指标的局限性;(2)提出新的语义感知评估方法;(3)验证其对语义和语调变化的鲁棒性;(4)探索跨数据集泛化挑战。在PHOENIX-14T和CSL-Daily数据集上,SiLVERScore实现接近完美的判别能力(ROC AUC = 0.99,重叠<7%),显著优于传统指标。

原文摘要 · Abstract (English)

Evaluating sign language generation is often done through back-translation, where generated signs are first recognized back to text and then compared to a reference using text-based metrics. However, this two-step evaluation pipeline introduces ambiguity: it not only fails to capture the multimodal nature of sign language-such as facial expressions, spatial grammar, and prosody-but also makes it hard to pinpoint whether evaluation errors come from sign generation model or the translation system used to assess it. In this work, we propose SiLVERScore, a novel semantically-aware embedding-based evaluation metric that assesses sign language generation in a joint embedding space. Our contributions include: (1) identifying limitations of existing metrics, (2) introducing SiLVERScore for semantically-aware evaluation, (3) demonstrating its robustness to semantic and prosodic variations, and (4) exploring generalization challenges across datasets. On PHOENIX-14T and CSL-Daily datasets, SiLVERScore achieves near-perfect discrimination between correct and random pairs (ROC AUC = 0.99, overlap < 7%), substantially outperforming traditional metrics.

手语生成评估方法嵌入空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。