Siformer提升手语识别精度,通过真实手部姿态与局部特征隔离增强模型鲁棒性。
Siformer: Feature-isolated Transformer for Efficient Skeleton-based Sign Language Recognition
- 引入运动学手姿修正,使手部骨骼表示更贴近真实动作
- 在WLASL100上达86.50%准确率,较前序方法提升2.39%
- 适配不同复杂度手势,兼顾效率与精度,适合实际部署
手语识别(SLR)旨在自动解析视频中的手语词义。由于手语包含快速且复杂的动作,如手势、身体姿势和面部表情,该任务在计算机视觉中极具挑战性。近年来,基于骨骼的动作识别因能独立处理个体与背景差异而受到关注。然而现有骨架式SLR方法存在三大局限:1)忽视真实手部姿态,多数研究使用非真实的骨骼表示训练模型;2)假设训练与推理阶段数据完整,且集体捕捉各身体部位间复杂关系;3)对所有手语词一视同仁,未考虑其骨骼表示复杂度的差异。为此,本文提出一种运动学手姿修正方法以增强手部姿态真实性;针对缺失数据问题,设计特征隔离机制,独立捕获各局部时空上下文,提升模型鲁棒性;此外,提出输入自适应推理策略,动态优化计算效率与准确性。实验表明,本方法在WLASL100和LSA64上均达到新SOTA:WLASL100上顶1准确率达86.50%,相对前序SOTA提升2.39%;LSA64上达99.84%。
原文摘要 · Abstract (English)
Sign language recognition (SLR) refers to interpreting sign language glosses from given videos automatically. This research area presents a complex challenge in computer vision because of the rapid and intricate movements inherent in sign languages, which encompass hand gestures, body postures, and even facial expressions. Recently, skeleton-based action recognition has attracted increasing attention due to its ability to handle variations in subjects and backgrounds independently. However, current skeleton-based SLR methods exhibit three limitations: 1) they often neglect the importance of realistic hand poses, where most studies train SLR models on non-realistic skeletal representations; 2) they tend to assume complete data availability in both training or inference phases, and capture intricate relationships among different body parts collectively; 3) these methods treat all sign glosses uniformly, failing to account for differences in complexity levels regarding skeletal representations. To enhance the realism of hand skeletal representations, we present a kinematic hand pose rectification method for enforcing constraints. Mitigating the impact of missing data, we propose a feature-isolated mechanism to focus on capturing local spatial-temporal context. This method captures the context concurrently and independently from individual features, thus enhancing the robustness of the SLR model. Additionally, to adapt to varying complexity levels of sign glosses, we develop an input-adaptive inference approach to optimise computational efficiency and accuracy. Experimental results demonstrate the effectiveness of our approach, as evidenced by achieving a new state-of-the-art (SOTA) performance on WLASL100 and LSA64. For WLASL100, we achieve a top-1 accuracy of 86.50\%, marking a relative improvement of 2.39% over the previous SOTA. For LSA64, we achieve a top-1 accuracy of 99.84%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。