构建双视角中文手语数据集,提升手语识别准确率
Dual-view Spatio-Temporal Feature Fusion with CNN-Transformer Hybrid Network for Chinese Isolated Sign Language Recognition
- 设计双视角融合网络,用CNN与Transformer结合提取特征
- 新数据集覆盖全部国家标准手语词汇,含13.4万段视频
- 适合做手语识别、多视角视觉分析的研究者使用
近年来,随着手语数据集的增多以及深度神经网络的发展,孤立手语识别(ISLR)取得了显著进展。然而,实际应用仍面临挑战:现有数据集未覆盖完整手语词汇,且大多仅提供单视角RGB视频,难以处理手部遮挡问题。为此,本文提出一个名为NationalCSL-DP的双视角手语数据集,全面涵盖中国国家标准手语词汇。该数据集包含10位手语者录制的134,140段视频,分别从正面和左侧两个垂直视角获取。同时,提出一种基于CNN-Transformer的混合网络作为强基线,并设计了一种简单但高效的融合策略进行预测。大量实验验证了数据集与基线的有效性。结果表明,所提融合策略能显著提升识别性能,但序列到序列模型无论采用早期融合还是晚期融合,均难以有效学习双视角视频间的互补特征。
原文摘要 · Abstract (English)
Due to the emergence of many sign language datasets, isolated sign language recognition (ISLR) has made significant progress in recent years. In addition, the development of various advanced deep neural networks is another reason for this breakthrough. However, challenges remain in applying the technique in the real world. First, existing sign language datasets do not cover the whole sign vocabulary. Second, most of the sign language datasets provide only single view RGB videos, which makes it difficult to handle hand occlusions when performing ISLR. To fill this gap, this paper presents a dual-view sign language dataset for ISLR named NationalCSL-DP, which fully covers the Chinese national sign language vocabulary. The dataset consists of 134140 sign videos recorded by ten signers with respect to two vertical views, namely, the front side and the left side. Furthermore, a CNN transformer network is also proposed as a strong baseline and an extremely simple but effective fusion strategy for prediction. Extensive experiments were conducted to prove the effectiveness of the datasets as well as the baseline. The results show that the proposed fusion strategy can significantly increase the performance of the ISLR, but it is not easy for the sequence-to-sequence model, regardless of whether the early-fusion or late-fusion strategy is applied, to learn the complementary features from the sign videos of two vertical views.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。