arXiv:2505.17090cs.CV2025-05被引 7

首个带情绪标注的美国手语多模态数据集,助力聋人沟通理解。

EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language

  • 采集200段美式手语视频,由3位聋人专业译员标注情感与语气。
  • 包含手部动作与面部表情的双模态情绪线索,支持细粒度识别。
  • 适合研究手语情感识别、无障碍沟通与多模态模型评估者使用。

与口语中韵律特征表达情感已有深入研究不同,手语中的情绪表现机制仍不清晰,导致在重要场景下存在沟通障碍。手语的独特之处在于面部表情和手势动作同时承担语法与情感双重功能。为填补这一空白,我们推出了EmoSign,这是首个包含200段美式手语(ASL)视频的情感与情绪标签的多模态数据集。我们还收集了关于情绪线索的开放式描述。所有标注均由3位具有专业口译经验的聋人ASL使用者完成。除标注数据外,我们提供了情绪与语气分类的基线模型。该数据集不仅弥补了现有手语研究的关键空白,也为手语多模态情绪识别模型的能力评估建立了新基准。数据集已公开于https://huggingface.co/datasets/catfang/emosign。

原文摘要 · Abstract (English)

Unlike spoken languages where the use of prosodic features to convey emotion is well studied, indicators of emotion in sign language remain poorly understood, creating communication barriers in critical settings. Sign languages present unique challenges as facial expressions and hand movements simultaneously serve both grammatical and emotional functions. To address this gap, we introduce EmoSign, the first sign video dataset containing sentiment and emotion labels for 200 American Sign Language (ASL) videos. We also collect open-ended descriptions of emotion cues. Annotations were done by 3 Deaf ASL signers with professional interpretation experience. Alongside the annotations, we include baseline models for sentiment and emotion classification. This dataset not only addresses a critical gap in existing sign language research but also establishes a new benchmark for understanding model capabilities in multimodal emotion recognition for sign languages. The dataset is made available at https://huggingface.co/datasets/catfang/emosign.

手语识别多模态情感分析无障碍

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。