arXiv:2604.27232cs.CL2026-04中稿 · CVPR

用最小翻译对分析手语模型,发现其依赖手势却忽略面部表情。

Targeted Linguistic Analysis of Sign Language Models with Minimal Translation Pairs

  • 构建了美国手语最小翻译对数据集,用于精准测试模型语言理解能力。
  • 模型在多数语言现象上表现高于随机水平,但严重依赖手部动作。
  • 揭示了手语模型对非手动线索(如面部表情)的忽视,适合手语技术研究者参考。

手语模型长期以来落后于语音和文本模型。尽管近期在手语翻译和孤立手势识别任务上取得显著进展,但现有模型对多种手语语言现象的理解程度仍不明确,且其对多发音器(手、上身、面部)线索的利用情况尚不清楚。本文提出一个针对美国手语的新基准数据集——ASL 最小翻译对(ASL-MTP),按手语现象分类并设计对应最小翻译对,以开展精细化语言分析。作为案例研究,我们使用 ASL-MTP 分析了一个最先进的手语到英语翻译模型。通过在训练和推理阶段剔除不同输入线索,评估模型在各类现象上的表现。结果表明,尽管模型在多数现象上表现优于随机水平,但高度依赖手部线索,常忽略关键的非手动线索。

原文摘要 · Abstract (English)

Models of sign language have historically lagged behind those for spoken language (text and speech). Recent work has greatly improved their performance on tasks like sign language translation and isolated sign recognition. However, it remains unclear to what extent existing models capture various linguistic phenomena of sign language, and how well they use cues from the multiple articulators used in sign language (hands, upper body, face). We introduce a new benchmark dataset for American Sign Language, ASL Minimal Translation Pairs (ASL-MTP), divided into multiple types of sign language phenomena and corresponding minimal pairs of translations, for performing such linguistic analyses. As a case study, we use ASL-MTP to analyze a state-of-the-art ASL-to-English translation model. We conduct a targeted analysis of the model by ablating various input cues during training and inference and evaluating on the phenomena in ASL-MTP. Our results show that, while the model performs above chance level on most of the phenomena, it relies strongly on manual cues while often missing crucial non-manual cues.

手语模型语言分析多模态非手动线索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。