arXiv:2511.05772cs.CVcs.AI2025-11

用骨骼数据识别手语,结合图网络与循环网络提升准确率。

Sign language recognition from skeletal data using graph and recurrent neural networks

  • 构建图-循环神经网络,同时捕捉手势空间结构和时间动态
  • 在AUTSL数据集上达到高识别准确率,验证方法有效性
  • 适合关注动作识别与手语理解的研究者

本文提出一种基于骨骼姿态数据的孤立手语手势识别方法,从视频序列中提取骨架信息。采用图-GRU时序网络建模帧间空间与时间依赖关系,实现精准分类。模型在安卡拉大学土耳其手语(AUTSL)数据集上训练与评估,实验结果表明,将基于图的空间表示与时序建模相结合具有显著效果,为手语识别提供了可扩展的框架。该方法凸显了基于姿态的手语理解潜力。

原文摘要 · Abstract (English)

This work presents an approach for recognizing isolated sign language gestures using skeleton-based pose data extracted from video sequences. A Graph-GRU temporal network is proposed to model both spatial and temporal dependencies between frames, enabling accurate classification. The model is trained and evaluated on the AUTSL (Ankara university Turkish sign language) dataset, achieving high accuracy. Experimental results demonstrate the effectiveness of integrating graph-based spatial representations with temporal modeling, providing a scalable framework for sign language recognition. The results of this approach highlight the potential of pose-driven methods for sign language understanding.

手语识别骨骼数据图网络序列建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。