arXiv:2506.21592cs.CLcs.CV2025-06被引 1

用骨架序列独立处理坐标,实现高效高精度手语识别

SignBart -- New approach with the skeleton sequence for Isolated Sign language Recognition

  • 分离编码坐标x/y,通过交叉注意力保持关联性
  • 仅74.9万参数,达96.04%准确率,超越百万参数模型
  • 适用于聋哑人群体的无障碍交流工具开发

手语识别对听力障碍者打破沟通障碍至关重要。以往方法在效率与精度间难以兼顾:RNN、LSTM和GCN存在梯度消失与高计算成本问题;尽管性能提升,基于Transformer的方法仍未普及。本研究提出SignBart新方法,突破传统模型将骨架序列的x、y坐标视为不可分的局限,采用BART架构的编码器-解码器结构,分别独立编码x、y坐标,同时通过交叉注意力维持二者关联。模型仅含749,888个参数,在LSA-64数据集上达到96.04%准确率,显著优于参数超百万的现有模型。在WLASL和ASL-Citizen数据集上也展现优异泛化能力。消融实验表明,坐标投影、归一化及多骨架组件使用对性能提升至关重要。该方法为听障群体提供可靠有效的手语识别方案,具备增强无障碍工具的潜力。

原文摘要 · Abstract (English)

Sign language recognition is crucial for individuals with hearing impairments to break communication barriers. However, previous approaches have had to choose between efficiency and accuracy. Such as RNNs, LSTMs, and GCNs, had problems with vanishing gradients and high computational costs. Despite improving performance, transformer-based methods were not commonly used. This study presents a new novel SLR approach that overcomes the challenge of independently extracting meaningful information from the x and y coordinates of skeleton sequences, which traditional models often treat as inseparable. By utilizing an encoder-decoder of BART architecture, the model independently encodes the x and y coordinates, while Cross-Attention ensures their interrelation is maintained. With only 749,888 parameters, the model achieves 96.04% accuracy on the LSA-64 dataset, significantly outperforming previous models with over one million parameters. The model also demonstrates excellent performance and generalization across WLASL and ASL-Citizen datasets. Ablation studies underscore the importance of coordinate projection, normalization, and using multiple skeleton components for boosting model efficacy. This study offers a reliable and effective approach for sign language recognition, with strong potential for enhancing accessibility tools for the deaf and hard of hearing.

手语识别骨架序列BART无障碍

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。