arXiv:2410.00681cs.CVcs.AI2024-10被引 10

用Transformer模型提升阿拉伯手语识别准确率,最高达99.6%

Advanced Arabic Alphabet Sign Language Recognition Using Transfer Learning and Transformer Models

  • 结合迁移学习与ViT/Swin等视觉变压器模型,捕捉手语动作特征
  • 在ArSL2018和AASL数据集上分别达到99.6%和99.43%准确率
  • 为阿拉伯语聋哑人群提供更高效沟通工具,推动社会包容性

本文提出一种基于深度学习的阿拉伯字母手语识别方法,融合迁移学习与基于Transformer的模型。研究对比了ResNet50、MobileNetV2、EfficientNetB7等主流CNN架构,以及Google ViT和Microsoft Swin Transformer等最新视觉变压器模型在两个公开数据集ArSL2018和AASL上的表现。这些预训练模型在数据集上进行微调,以捕捉阿拉伯手语动作的独特特征。实验结果表明,所提方法在ArSL2018和AASL数据集上的识别准确率分别达到99.6%和99.43%,显著优于此前报道的最先进方法。该性能为阿拉伯语聋哑人群提供了更高效、无障碍的交流途径,有助于构建更具包容性的社会。

原文摘要 · Abstract (English)

This paper presents an Arabic Alphabet Sign Language recognition approach, using deep learning methods in conjunction with transfer learning and transformer-based models. We study the performance of the different variants on two publicly available datasets, namely ArSL2018 and AASL. This task will make full use of state-of-the-art CNN architectures like ResNet50, MobileNetV2, and EfficientNetB7, and the latest transformer models such as Google ViT and Microsoft Swin Transformer. These pre-trained models have been fine-tuned on the above datasets in an attempt to capture some unique features of Arabic sign language motions. Experimental results present evidence that the suggested methodology can receive a high recognition accuracy, by up to 99.6\% and 99.43\% on ArSL2018 and AASL, respectively. That is far beyond the previously reported state-of-the-art approaches. This performance opens up even more avenues for communication that may be more accessible to Arabic-speaking deaf and hard-of-hearing, and thus encourages an inclusive society.

手语识别Transformer阿拉伯语无障碍

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。