用视频Transformer提升手语识别准确率,突破传统模型瓶颈。
Breaking the Barriers: Video Vision Transformers for Word-Level Sign Language Recognition
- 采用ViViT架构,通过自注意力捕捉手势时空全局关系。
- 在WLASL100数据集上达到75.58%的准确率,优于CNN的65.89%。
- 适合关注手语识别、无障碍交互与Transformer应用的研究者。
手语是聋哑人群体沟通的核心方式,通过手势、面部表情和身体动作传递细腻信息。然而,听人普遍缺乏手语能力,导致交流障碍依然存在。自动手语识别(SLR)尤其在动态词级任务中面临挑战,需同时建模时空依赖性。尽管卷积神经网络(CNN)在该任务中有一定表现,但计算成本高且难以捕捉视频序列间的全局时间依赖。为此,本文提出基于视频视觉变压器(ViViT)的美国手语(ASL)词级识别模型。该模型利用自注意力机制,在空间与时间维度上高效建模长程依赖,显著提升识别性能。在WLASL100数据集上,VideoMAE模型达到75.58%的Top-1准确率,相较传统CNN的65.89%有明显提升。研究证实,基于Transformer的架构在手语识别中具有巨大潜力,可有效缓解沟通障碍,推动聋哑群体的社会融入。
原文摘要 · Abstract (English)
Sign language is a fundamental means of communication for the deaf and hard-of-hearing (DHH) community, enabling nuanced expression through gestures, facial expressions, and body movements. Despite its critical role in facilitating interaction within the DHH population, significant barriers persist due to the limited fluency in sign language among the hearing population. Overcoming this communication gap through automatic sign language recognition (SLR) remains a challenge, particularly at a dynamic word-level, where temporal and spatial dependencies must be effectively recognized. While Convolutional Neural Networks (CNNs) have shown potential in SLR, they are computationally intensive and have difficulties in capturing global temporal dependencies between video sequences. To address these limitations, we propose a Video Vision Transformer (ViViT) model for word-level American Sign Language (ASL) recognition. Transformer models make use of self-attention mechanisms to effectively capture global relationships across spatial and temporal dimensions, which makes them suitable for complex gesture recognition tasks. The VideoMAE model achieves a Top-1 accuracy of 75.58% on the WLASL100 dataset, highlighting its strong performance compared to traditional CNNs with 65.89%. Our study demonstrates that transformer-based architectures have great potential to advance SLR, overcome communication barriers and promote the inclusion of DHH individuals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。