arXiv:2509.03467cs.CLcs.AI2025-09被引 3

首个连续沙特手语数据集+高效识别模型,助力听障者沟通平等

Continuous Saudi Sign Language Recognition: A Vision Transformer Approach

  • 构建首个连续沙特手语数据集KAU-CSSL,支持完整句子识别
  • 提出融合ResNet与Transformer的模型,在识别人模式下达99.02%准确率
  • 为阿拉伯语手语研究提供基础,适合残障技术与AI多模态方向研究者

手语是聋哑人士重要的交流方式,有助于其融入社会。在沙特阿拉伯,超过84,000人依赖沙特手语(SSL)作为主要交流手段,但公众认知不足导致其在教育和职业机会上面临不平等,加剧社会排斥。尽管已有技术辅助聋哑人士沟通,但针对阿拉伯语手语(如SSL)的精准、可靠翻译技术仍严重缺乏。现有主流方法多集中于非阿拉伯语手语,阿拉伯语手语资源稀缺,且多数数据集仅包含孤立手势而非连续语句。为此,本文首次构建连续沙特手语数据集KAU-CSSL,聚焦完整句子以推动后续研究。同时提出基于Transformer的识别模型:采用预训练ResNet-18提取空间特征,结合双向LSTM与Transformer编码器建模时间依赖性,在签名人依赖模式下达到99.02%准确率,签名人独立模式下达77.71%。该成果不仅提升SSL沟通工具性能,也为手语识别领域做出重要贡献。

原文摘要 · Abstract (English)

Sign language (SL) is an essential communication form for hearing-impaired and deaf people, enabling engagement within the broader society. Despite its significance, limited public awareness of SL often leads to inequitable access to educational and professional opportunities, thereby contributing to social exclusion, particularly in Saudi Arabia, where over 84,000 individuals depend on Saudi Sign Language (SSL) as their primary form of communication. Although certain technological approaches have helped to improve communication for individuals with hearing impairments, there continues to be an urgent requirement for more precise and dependable translation techniques, especially for Arabic sign language variants like SSL. Most state-of-the-art solutions have primarily focused on non-Arabic sign languages, resulting in a considerable absence of resources dedicated to Arabic sign language, specifically SSL. The complexity of the Arabic language and the prevalence of isolated sign language datasets that concentrate on individual words instead of continuous speech contribute to this issue. To address this gap, our research represents an important step in developing SSL resources. To address this, we introduce the first continuous Saudi Sign Language dataset called KAU-CSSL, focusing on complete sentences to facilitate further research and enable sophisticated recognition systems for SSL recognition and translation. Additionally, we propose a transformer-based model, utilizing a pretrained ResNet-18 for spatial feature extraction and a Transformer Encoder with Bidirectional LSTM for temporal dependencies, achieving 99.02\% accuracy at signer dependent mode and 77.71\% accuracy at signer independent mode. This development leads the way to not only improving communication tools for the SSL community but also making a substantial contribution to the wider field of sign language.

手语识别视觉变换器沙特手语连续识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。