arXiv:2510.11243cs.CVcs.AI2025-10

首个尼泊尔手语数据集及深度学习识别研究

Nepali Sign Language Characters Recognition: Dataset Development and Deep Learning Approaches

  • 构建36类手语动作数据集,每类1500样本
  • 移动端模型达90.45%准确率,验证小样本有效性
  • 为残障人群沟通技术提供基础,适合语音识别研究者

手语是听障与言语障碍者的重要交流方式,但尼泊尔手语(NSL)等非主流手语的数字语料资源仍严重匮乏。本研究首次构建了首个NSL基准数据集,包含36个手势类别,每类1500个样本,旨在捕捉该语言的结构与视觉特征。为评估识别性能,我们对MobileNetV2和ResNet50进行微调,分类准确率分别达到90.45%和88.78%。结果表明卷积神经网络在低资源手语识别任务中具有有效性。据我们所知,这是首次系统性构建基准数据集并评估深度学习方法用于NSL识别的工作,凸显了迁移学习与微调在推进未充分研究手语研究中的潜力。

原文摘要 · Abstract (English)

Sign languages serve as essential communication systems for individuals with hearing and speech impairments. However, digital linguistic dataset resources for underrepresented sign languages, such as Nepali Sign Language (NSL), remain scarce. This study introduces the first benchmark dataset for NSL, consisting of 36 gesture classes with 1,500 samples per class, designed to capture the structural and visual features of the language. To evaluate recognition performance, we fine-tuned MobileNetV2 and ResNet50 architectures on the dataset, achieving classification accuracies of 90.45% and 88.78%, respectively. These findings demonstrate the effectiveness of convolutional neural networks in sign recognition tasks, particularly within low-resource settings. To the best of our knowledge, this work represents the first systematic effort to construct a benchmark dataset and assess deep learning approaches for NSL recognition, highlighting the potential of transfer learning and fine-tuning for advancing research in underexplored sign languages.

手语识别小样本数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。