arXiv:2411.02768cs.CV2024-11被引 4

构建首个泰语单阶段指拼手势数据集,助力手语识别研究

One-Stage-TFS: Thai One-Stage Fingerspelling Dataset for Fingerspelling Recognition Frameworks

  • 收集7200张泰语单阶段辅音手势图像,涵盖有无手语经验者
  • 包含15种手势,支持端到端识别框架的开发与验证
  • 适合手语识别、计算机视觉方向的研究者使用

泰国单阶段指拼(One-Stage-TFS)数据集是一个全面的手势识别资源,专门用于推进泰语手语识别研究。该数据集包含7,200张图像,记录了泰国玛哈萨拉堪皇家大学本科生演示的15种单阶段辅音手势。参与者包括精通泰语手语的特殊教育专业学生以及无手语基础的其他专业学生。图像采集时间为2021年7月至12月,使用DSLR相机拍摄,背景包括简单和复杂场景。该数据集在手势检测与识别方面具有挑战性,为开发新型端到端识别框架提供了可能。研究人员可利用此数据集探索YOLO、EfficientDet、RetinaNet、Detectron等深度学习方法进行手部检测,再结合CNN、Transformer及自适应特征融合网络进行特征提取与识别。数据集可通过Mendeley Data库获取,广泛适用于深度学习、计算机视觉与模式识别等领域,推动相关技术创新与探索。

原文摘要 · Abstract (English)

The Thai One-Stage Fingerspelling (One-Stage-TFS) dataset is a comprehensive resource designed to advance research in hand gesture recognition, explicitly focusing on the recognition of Thai sign language. This dataset comprises 7,200 images capturing 15 one-stage consonant gestures performed by undergraduate students from Rajabhat Maha Sarakham University, Thailand. The contributors include both expert students from the Special Education Department with proficiency in Thai sign language and students from other departments without prior sign language experience. Images were collected between July and December 2021 using a DSLR camera, with contributors demonstrating hand gestures against both simple and complex backgrounds. The One-Stage-TFS dataset presents challenges in detecting and recognizing hand gestures, offering opportunities to develop novel end-to-end recognition frameworks. Researchers can utilize this dataset to explore deep learning methods, such as YOLO, EfficientDet, RetinaNet, and Detectron, for hand detection, followed by feature extraction and recognition using techniques like convolutional neural networks, transformers, and adaptive feature fusion networks. The dataset is accessible via the Mendeley Data repository and supports a wide range of applications in computer science, including deep learning, computer vision, and pattern recognition, thereby encouraging further innovation and exploration in these fields.

手语识别数据集计算机视觉泰语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。