arXiv:2512.10562cs.CV2025-12

用少量数据实现高效手语识别,解决罕见手势识别难题。

Data-Efficient American Sign Language Recognition via Few-Shot Prototypical Networks

  • 基于骨架信息构建原型网络,通过动态原型比对实现少样本分类。
  • 在WLASL数据集上达43.75%准确率,较传统方法提升超13%。
  • 零样本迁移能力突出,无需微调即可在新数据集上保持近30%准确率。

孤立手语识别(ISLR)对弥合聋哑人群与听力者之间的沟通鸿沟至关重要。然而,由于手势词汇量庞大且数据稀缺,收集足够样本极其困难。标准分类方法在此情境下易过拟合常见类别,难以泛化到稀有手势。为此,我们提出一种基于骨架编码器的少样本原型网络框架。不同于固定决策边界,该方法通过任务式训练学习语义度量空间,依据手势与动态类别原型的接近程度进行分类。结合时空图卷积网络(ST-GCN)与新型多尺度时序聚合(MSTA)模块,有效捕捉快速与流畅的动作动态。在WLASL数据集上的实验表明,该模型在测试集上达到43.75%的Top-1和77.10%的Top-5准确率,显著优于共享相同骨干架构的标准分类基线(提升超13%)。此外,模型具备强零样本泛化能力,在未见的SignASL数据集上无需微调即达近30%准确率,为有限数据下大规模手语识别提供可扩展路径。

原文摘要 · Abstract (English)

Isolated Sign Language Recognition (ISLR) is critical for bridging the communication gap between the Deaf and Hard-of-Hearing (DHH) community and the hearing world. However, robust ISLR is fundamentally constrained by data scarcity and the long-tail distribution of sign vocabulary, where gathering sufficient examples for thousands of unique signs is prohibitively expensive. Standard classification approaches struggle under these conditions, often overfitting to frequent classes while failing to generalize to rare ones. To address this bottleneck, we propose a Few-Shot Prototypical Network framework adapted for a skeleton based encoder. Unlike traditional classifiers that learn fixed decision boundaries, our approach utilizes episodic training to learn a semantic metric space where signs are classified based on their proximity to dynamic class prototypes. We integrate a Spatiotemporal Graph Convolutional Network (ST-GCN) with a novel Multi-Scale Temporal Aggregation (MSTA) module to capture both rapid and fluid motion dynamics. Experimental results on the WLASL dataset demonstrate the superiority of this metric learning paradigm: our model achieves 43.75% Top-1 and 77.10% Top-5 accuracy on the test set. Crucially, this outperforms a standard classification baseline sharing the identical backbone architecture by over 13%, proving that the prototypical training strategy effectively outperforms in a data scarce situation where standard classification fails. Furthermore, the model exhibits strong zero-shot generalization, achieving nearly 30% accuracy on the unseen SignASL dataset without fine-tuning, offering a scalable pathway for recognizing extensive sign vocabularies with limited data.

手语识别少样本学习原型网络骨骼动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。