构建首个基于HamNoSys的精细手形识别数据集,支持手语技术开发。
A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language

- 用HamNoSys标准定义160类手形,采集15人共14.4万张图像。
- 跨参与者识别准确率下降明显,凸显泛化挑战。
- 提供可复现基准,适合手语识别与无障碍技术研究者。
细粒度手形识别有助于手语的计算转录、识别与翻译,但目前缺乏广义、音位定义的视觉库及考虑签名人差异的评估体系。本文基于通用的汉堡记号系统(HamNoSys),构建了一个包含160种手形类别、由15名参与者采集的14.4万张RGB图像的平衡数据集。评估了ResNet-18、ViT-B/16等外观模型,以及基于手部关键点的图卷积网络和XGBoost模型。采用类内分组的受试者相关划分和15折留一参与者外(LOSO)协议进行验证。同时在LSWH100和ASL Fingerspelling Dataset A上进行了外部对比。结果表明,受试者相关基准提供了可复现的性能参考,而LOSO评估揭示了跨参与者识别显著下降的问题。在ASL Fingerspelling Dataset A上,平均LOSO top-1准确率为82.20%至87.40%。本工作提供的采集、整理与评估流程为细粒度孤立手形研究提供了可复现资源,推动更普惠的手语技术发展。
原文摘要 · Abstract (English)
Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, phonetically defined visual inventories with signer-aware evaluation remain limited. This work introduces a benchmark grounded in the language-independent Hamburg Notation System (HamNoSys). Methods: A balanced dataset of 144,000 RGB images was collected from 15 participants for 160 handshape classes defined by the official HamNoSys 4 Handshapes Chart. ResNet-18 and ViT-B/16 were evaluated as appearance-based models, while a graph convolutional network and XGBoost were evaluated from hand landmarks. Both a class-stratified subject-dependent split and a 15-fold leave-one-subject-out (LOSO) protocol were used. The same model families were additionally assessed on LSWH100 and ASL Fingerspelling Dataset A for external context. Results: The subject-dependent benchmarks established reproducible reference performance across all four model families, whereas LOSO evaluation exposed a substantial reduction when recognition was required to generalise to unseen participants. On ASL Fingerspelling Dataset A, mean LOSO top-1 accuracy ranged from 82.20% to 87.40%. Conclusion: The documented acquisition, curation, and complementary evaluation protocols pro-vide a reproducible resource for fine-grained isolated-handshape research and for developing more accessible sign-language technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。