arXiv:2512.15376cs.CVcs.AI2025-12被引 1

用跨语言数据缓解手语情绪识别数据少难题,提升识别效果。

Emotion Recognition in Signers

  • 利用口语情绪文本缓解手语数据稀缺问题
  • 时间片段选择对识别效果影响显著,提升准确率
  • 融合手势运动信息,优于现有大模型基线

手语情绪识别面临两大挑战:语法性与情感性面部表情重叠,以及用于模型训练的数据稀缺。本文在跨语言场景下,基于新构建的eJSL数据集(含78个表达式、7种情绪状态,共1,092段视频)和大型英国手语数据集BOBSL(带字幕),提出解决方案。实验证明:1)口语中的文本情绪识别可缓解手语数据不足;2)时间片段选择对性能有显著影响;3)引入手势运动特征能有效提升情绪识别效果。最终建立的基线模型优于现有口语大模型。

原文摘要 · Abstract (English)

Recognition of signers' emotions suffers from one theoretical challenge and one practical challenge, namely, the overlap between grammatical and affective facial expressions and the scarcity of data for model training. This paper addresses these two challenges in a cross-lingual setting using our eJSL dataset, a new benchmark dataset for emotion recognition in Japanese Sign Language signers, and BOBSL, a large British Sign Language dataset with subtitles. In eJSL, two signers expressed 78 distinct utterances with each of seven different emotional states, resulting in 1,092 video clips. We empirically demonstrate that 1) textual emotion recognition in spoken language mitigates data scarcity in sign language, 2) temporal segment selection has a significant impact, and 3) incorporating hand motion enhances emotion recognition in signers. Finally we establish a stronger baseline than spoken language LLMs.

情绪识别手语理解跨语言数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。