arXiv:2608.10588cs.CVcs.AI2026-08

构建首个基于HamNoSys的精细手形识别数据集,支持手语技术开发。

A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language

论文配图:A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language
图 1 · 摘自论文原文
  • 用HamNoSys标准定义160类手形,采集15人共14.4万张图像。
  • 跨参与者识别准确率下降明显,凸显泛化挑战。
  • 提供可复现基准,适合手语识别与无障碍技术研究者。

细粒度手形识别有助于手语的计算转录、识别与翻译,但目前缺乏广义、音位定义的视觉库及考虑签名人差异的评估体系。本文基于通用的汉堡记号系统(HamNoSys),构建了一个包含160种手形类别、由15名参与者采集的14.4万张RGB图像的平衡数据集。评估了ResNet-18、ViT-B/16等外观模型,以及基于手部关键点的图卷积网络和XGBoost模型。采用类内分组的受试者相关划分和15折留一参与者外(LOSO)协议进行验证。同时在LSWH100和ASL Fingerspelling Dataset A上进行了外部对比。结果表明,受试者相关基准提供了可复现的性能参考,而LOSO评估揭示了跨参与者识别显著下降的问题。在ASL Fingerspelling Dataset A上,平均LOSO top-1准确率为82.20%至87.40%。本工作提供的采集、整理与评估流程为细粒度孤立手形研究提供了可复现资源,推动更普惠的手语技术发展。

原文摘要 · Abstract (English)

Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, phonetically defined visual inventories with signer-aware evaluation remain limited. This work introduces a benchmark grounded in the language-independent Hamburg Notation System (HamNoSys). Methods: A balanced dataset of 144,000 RGB images was collected from 15 participants for 160 handshape classes defined by the official HamNoSys 4 Handshapes Chart. ResNet-18 and ViT-B/16 were evaluated as appearance-based models, while a graph convolutional network and XGBoost were evaluated from hand landmarks. Both a class-stratified subject-dependent split and a 15-fold leave-one-subject-out (LOSO) protocol were used. The same model families were additionally assessed on LSWH100 and ASL Fingerspelling Dataset A for external context. Results: The subject-dependent benchmarks established reproducible reference performance across all four model families, whereas LOSO evaluation exposed a substantial reduction when recognition was required to generalise to unseen participants. On ASL Fingerspelling Dataset A, mean LOSO top-1 accuracy ranged from 82.20% to 87.40%. Conclusion: The documented acquisition, curation, and complementary evaluation protocols pro-vide a reproducible resource for fine-grained isolated-handshape research and for developing more accessible sign-language technologies.

手语识别细粒度识别数据集多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。