用图神经网络提升手语手势形状识别,准确率超基线一倍。
Improving Handshape Representations for Sign Language Processing: A Graph Neural Network Approach
- 构建基于解剖结构的图模型,分离手势动态与静态形状特征。
- 在37类手势形状上达到46%准确率,显著优于25%的基线方法。
- 首次建立结构化手势形状识别基准,适合手语技术与语言学研究者。
手势形状在手语中具有基础音位作用,美国手语包含约50种不同形状。然而,现有计算方法很少显式建模手势形状,限制了识别准确率和语言分析能力。本文提出一种新型图神经网络,将时间动态与静态手势形状配置分离。该方法结合解剖学启发的图结构与对比学习,解决手势形状识别中的细微类别差异和时间变化难题。我们建立了首个针对签名序列中结构化手势形状识别的基准测试,在37个手势形状类别上实现46%的准确率(基线方法为25%)。
原文摘要 · Abstract (English)
Handshapes serve a fundamental phonological role in signed languages, with American Sign Language employing approximately 50 distinct shapes. However,computational approaches rarely model handshapes explicitly, limiting both recognition accuracy and linguistic analysis.We introduce a novel graph neural network that separates temporal dynamics from static handshape configurations. Our approach combines anatomically-informed graph structures with contrastive learning to address key challenges in handshape recognition, including subtle interclass distinctions and temporal variations. We establish the first benchmark for structured handshape recognition in signing sequences, achieving 46% accuracy across 37 handshape classes (with baseline methods achieving 25%).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。