构建首个大规模俄语手语数据集Logos,提升跨语言手语识别性能。
Logos as a Well-Tempered Pre-train for Sign Language Recognition
- 构建包含大量签名者与词汇的俄语手语数据集Logos。
- 基于Logos预训练模型在多个数据集上实现最优或竞争性准确率。
- 显式标注视觉相似手势组,提升模型作为视觉编码器的泛化能力。
本文研究孤立手语识别(ISLR)的两个关键问题:其一,尽管已有若干数据集,但各手语语言的数据量仍有限,导致跨语言模型训练困难;其二,外观相似的手势可能含义不同,造成标签歧义。为此,本文提出Logos——目前规模最大、签名人最多、词汇最丰富的俄语手语(RSL)数据集,也是最大的RSL数据集。实验表明,基于Logos预训练的模型可作为通用视觉编码器,适用于其他手语识别任务,包括少样本学习。通过多语言联合训练与多分类头策略,显著提升低资源数据集上的识别精度。数据集关键特性为显式标注视觉相似手势组,实验显示此举可有效提升下游任务的模型质量。基于此,本方法在WLASL数据集上超越现有最佳结果,在AUTSL数据集上取得具有竞争力的表现,仅使用单流RGB视频模型。代码、数据集及预训练模型均已公开。
原文摘要 · Abstract (English)
This paper examines two aspects of the isolated sign language recognition (ISLR) task. First, although a certain number of datasets is available, the data for individual sign languages is limited. It poses the challenge of cross-language ISLR model training, including transfer learning. Second, similar signs can have different semantic meanings. It leads to ambiguity in dataset labeling and raises the question of the best policy for annotating such signs. To address these issues, this study presents Logos, a novel Russian Sign Language (RSL) dataset, the most extensive available ISLR dataset by the number of signers, one of the most extensive datasets in size and vocabulary, and the largest RSL dataset. It is shown that a model, pre-trained on the Logos dataset can be used as a universal encoder for other language SLR tasks, including few-shot learning. We explore cross-language transfer learning approaches and find that joint training using multiple classification heads benefits accuracy for the target low-resource datasets the most. The key feature of the Logos dataset is explicitly annotated visually similar sign groups. We show that explicitly labeling visually similar signs improves trained model quality as a visual encoder for downstream tasks. Based on the proposed contributions, we outperform current state-of-the-art results for the WLASL dataset and get competitive results for the AUTSL dataset, with a single stream model processing solely RGB video. The source code, dataset, and pre-trained models are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。