arXiv:2605.23328cs.CL2026-05

构建首个手语对话情绪数据集,揭示现有模型在真实场景下的性能瓶颈。

Emotion Recognition in Sign Language Conversation

论文配图:Emotion Recognition in Sign Language Conversation
图 1 · 摘自论文原文
  • 提出eJSL Dialog数据集,包含480组对话共1920个视频样本。
  • 发现通用多模态模型在手语对话中表现显著下降,存在领域差距。
  • 强调需专用于手语的上下文感知视觉提取器,推动大规模预训练发展。

情感识别是情感计算的核心组成部分,而当前的手语情感数据集主要聚焦于孤立语句,缺乏对话上下文。仅在孤立语句上训练的模型在真实场景中表现下降,因其无法利用历史对话流。为解决这一结构性局限,我们引入手语视频分析中的对话情感识别(ERC)任务,并提出eJSL Dialog数据集。该数据集基于STUDIES语料库脚本构建,包含1,920个视频样本,组织成480个唯一对话。我们在该数据集上对从孤立视觉网络到多模态对话架构的多种模型进行了系统基准测试。结果表明,将通用多模态对话情感识别模型应用于手语时存在明显领域差距。这些发现凸显了对手语专用上下文感知视觉提取器的迫切需求,并指出构建更大规模对话数据集以支持大规模预训练是未来研究的必要方向。

原文摘要 · Abstract (English)

Emotion Recognition in Conversation is a core component of affective computing, while current sign language emotion datasets primarily focus on isolated sentences and lack conversational context. Models trained exclusively on these isolated utterances demonstrate degraded performance in real world scenarios because they cannot utilize historical dialogue flow. To address this structural limitation, we introduce the ERC task to sign language video analysis and propose the eJSL Dialog dataset. Constructed using the scripts from the STUDIES corpus, the dataset contains 1,920 video samples organized into 480 unique dialogues. We conduct systematic benchmarking on this dataset using models ranging from isolated visual networks to multimodal conversational architectures. The results reveal a domain gap when applying generic multimodal conversational emotion recognition models to sign language. These findings demonstrate the explicit need for context-aware visual extractors specific to sign language and indicate that constructing larger conversational datasets to support large-scale pre-training is a necessary next step for future research.

手语识别情感计算对话建模多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。