发现手语生成中真实手势难以保持,现有评估指标不可靠。
Conditional Collapse in Sign Language Production: A Diagnostic and a Scaling Argument

- 用三个独立维度评估手语生成质量:起始姿态、输出多样性、目标忠实度。
- 在How2Sign数据集上,忠实度始终未达标,而传统指标波动近两个数量级。
- 小规模单字数据集可实现高忠实度,说明句子级数据量是主要瓶颈。
手语生成(SLP)是从自然语言文本生成虚拟人物手语动作的任务。现有评估通常依赖运动空间弗雷歇距离(FID)和回译BLEU分数,但在这些指标显著提升时,生成的手语动作可能仍不忠实。本文提出在三个独立层面评估生成动作:(τ1) 起始姿态条件性,(τ2) 输出多样性,(τ3) 目标忠实度。通过冻结的运动自编码器(MoAE)的潜在表示计算成对距离比。我们在How2Sign数据集上评估了14个SLP模型检查点,包括重新实现的Neural Sign Actors(NSA),发现τ3忠实度从未达到,而FID变化接近两个数量级且与忠实度无关。在ASL3DWord单字词汇数据集上,τ3可实现良好表现,表明句子级配对数据规模是核心瓶颈。
原文摘要 · Abstract (English)
Sign Language Production (SLP) is the task of generating avatar sign language motion from natural language text. The quality of the generated motion is typically evaluated by a motion-space Fréchet distance (FID) and back-translation (BT) BLEU score on benchmarks such as How2Sign. Both metrics can improve substantially while the underlying generator fails to faithfully represent the sign language gestures. In this work we propose to evaluate the generated motion at three independent levels: ($\tau1$) initial-pose conditioning, ($\tau2$) output diversity, and ($\tau3$) target faithfulness. We compute these as pairwise-distance ratios using latent representations of a frozen motion autoencoder (MoAE). We evaluate 14 SLP model checkpoints on the How2Sign dataset, including a re-implemented Neural Sign Actors (NSA), and show that $\tau3$ faithfulness is never attained, while FID varies by nearly two orders of magnitude and is uncorrelated with faithfulness. We show that on the isolated gloss dataset ASL3DWord favorable $\tau3$ can be attained, hence isolating the size of the sentence-level paired-dataset as the bottleneck.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。