arXiv:2503.12451cs.CVcs.AI2025-03被引 2

首个伊朗手语识别数据集,支持手势与骨骼双重建模。

ISLR101: an Iranian Word-Level Sign Language Recognition Dataset

  • 构建首个公开的伊朗手语数据集,含101个手势视频。
  • 视觉模型达97.01%准确率,骨骼模型达94.02%。
  • 适合手语识别、残障辅助技术研究者使用。

手语识别需处理复杂多通道信息,如手形与动作,但常因数据不足而受限。为填补空白,我们推出ISLR101,首个公开的孤立伊朗手语识别数据集。该数据集包含4,614个视频,覆盖101个不同手势,由10名不同背景的签署者(3名聋人、2名手语翻译员、5名二语学习者)在多样背景下录制,分辨率为800x600像素,帧率为25帧/秒,并使用OpenPose提取骨骼姿态信息。我们建立了基于视觉外观和骨骼姿态的基准模型,在该数据集上分别达到97.01%和94.02%的测试准确率。同时发布训练、验证与测试划分,便于公平比较。

原文摘要 · Abstract (English)

Sign language recognition involves modeling complex multichannel information, such as hand shapes and movements while relying on sufficient sign language-specific data. However, sign languages are often under-resourced, posing a significant challenge for research and development in this field. To address this gap, we introduce ISLR101, the first publicly available Iranian Sign Language dataset for isolated sign language recognition. This comprehensive dataset includes 4,614 videos covering 101 distinct signs, recorded by 10 different signers (3 deaf individuals, 2 sign language interpreters, and 5 L2 learners) against varied backgrounds, with a resolution of 800x600 pixels and a frame rate of 25 frames per second. It also includes skeleton pose information extracted using OpenPose. We establish both a visual appearance-based and a skeleton-based framework as baseline models, thoroughly training and evaluating them on ISLR101. These models achieve 97.01% and 94.02% accuracy on the test set, respectively. Additionally, we publish the train, validation, and test splits to facilitate fair comparisons.

手语识别数据集伊朗手语骨骼建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。