arXiv:2409.01073cs.CVcs.AI2024-09AAAI被引 11

用大模型嵌入增强手语识别与翻译的上下文理解能力。

SCOPE: Sign Language Contextual Processing with Embedding from LLMs

  • 结合对话上下文,通过多模态编码提升手语词级识别精度。
  • 在Phoenix-2014T、CSL-Daily等数据集上达到最新最好效果。
  • 开源72小时中文手语对话数据集,适合残障科技与AI研究者。

手语是全球约7000万聋人使用的视觉语言,传递视觉与上下文信息。当前基于视觉的手语识别(SLR)与翻译(SLT)方法在对话场景中表现不佳,主要受限于数据集多样性不足及对上下文信息的忽视。为此,本文提出SCOPE(Sign language Contextual Processing with Embedding from LLMs),一种新颖的上下文感知型视觉手语识别与翻译框架。针对手语识别,利用多模态编码器融合对话上下文以提升词级识别;针对手语翻译,进一步通过引入先前对话上下文微调大型语言模型(LLM)。同时,本文构建了一个新数据集,包含72小时中国手语视频,涵盖多种真实对话场景。实验表明,该框架在多个数据集(包括Phoenix-2014T、CSL-Daily及自建的SCOPE数据集)上均达到领先性能。与聋人群体的问卷调研也验证了该方法在实际应用中的鲁棒性与有效性。相关数据集与代码将开源,以推动后续研究。

原文摘要 · Abstract (English)

Sign languages, used by around 70 million Deaf individuals globally, are visual languages that convey visual and contextual information. Current methods in vision-based sign language recognition (SLR) and translation (SLT) struggle with dialogue scenes due to limited dataset diversity and the neglect of contextually relevant information. To address these challenges, we introduce SCOPE (Sign language Contextual Processing with Embedding from LLMs), a novel context-aware vision-based SLR and SLT framework. For SLR, we utilize dialogue contexts through a multi-modal encoder to enhance gloss-level recognition. For subsequent SLT, we further fine-tune a Large Language Model (LLM) by incorporating prior conversational context. We also contribute a new sign language dataset that contains 72 hours of Chinese sign language videos in contextual dialogues across various scenarios. Experimental results demonstrate that our SCOPE framework achieves state-of-the-art performance on multiple datasets, including Phoenix-2014T, CSL-Daily, and our SCOPE dataset. Moreover, surveys conducted with participants from the Deaf community further validate the robustness and effectiveness of our approach in real-world applications. Both our dataset and code will be open-sourced to facilitate further research.

手语识别上下文建模大模型应用无障碍AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。