用伪标签提升手语识别,少标注也能更准
SSLR: A Semi-Supervised Learning Method for Isolated Sign Language Recognition
- 用姿态数据+Transformer,通过伪标签挖掘无标注数据
- 在WLASL-100上,标签少时性能优于全监督模型
- 适合标注成本高、数据稀缺的手语识别场景
听障人士主要依赖手语沟通。手语识别(SLR)系统旨在识别手势并翻译为口语。当前主要挑战是标注数据稀缺。为此,本文提出一种半监督学习方法(SSLR),采用伪标签策略对无标注样本进行标注。手势通过编码签手骨骼关节点的姿态信息表示,并作为Transformer主干网络的输入。为验证不同标注比例下的学习能力,实验在不同标签比例和类别数下展开。在WLASL-100数据集上,与全监督模型相比,该半监督模型在多数情况下以更少标注数据实现了更高性能。
原文摘要 · Abstract (English)
Sign language is the primary communication language for people with disabling hearing loss. Sign language recognition (SLR) systems aim to recognize sign gestures and translate them into spoken language. One of the main challenges in SLR is the scarcity of annotated datasets. To address this issue, we propose a semi-supervised learning (SSL) approach for SLR (SSLR), employing a pseudo-label method to annotate unlabeled samples. The sign gestures are represented using pose information that encodes the signer's skeletal joint points. This information is used as input for the Transformer backbone model utilized in the proposed approach. To demonstrate the learning capabilities of SSL across various labeled data sizes, several experiments were conducted using different percentages of labeled data with varying numbers of classes. The performance of the SSL approach was compared with a fully supervised learning-based model on the WLASL-100 dataset. The obtained results of the SSL model outperformed the supervised learning-based model with less labeled data in many cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。