多尺度注意力机制提升手部动作识别准确率
Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition
- 采用多尺度多头自注意力构建金字塔特征结构
- 在NVGesture和Briareo数据集上分别达88.22%和99.10%准确率
- 支持单模态与多模态融合,适合手势识别场景
由于签名者手部姿态、大小和形状的差异,动态手势识别是极具挑战的研究方向。本文提出一种用于动态手部手势识别的多尺度多头自注意力视频变换网络(MsMHA-VTN)。通过变换器多尺度头注意力模型提取分层多尺度特征。所提模型为每个注意力头配置不同注意力维度,实现多尺度注意力。此外,还评估了单模态与多模态融合下的识别性能。大量实验表明,所提MsMHA-VTN在NVGesture和Briareo数据集上分别达到88.22%和99.10%的整体准确率。
原文摘要 · Abstract (English)
Dynamic gesture recognition is one of the challenging research areas due to variations in pose, size, and shape of the signer's hand. In this letter, Multiscaled Multi-Head Attention Video Transformer Network (MsMHA-VTN) for dynamic hand gesture recognition is proposed. A pyramidal hierarchy of multiscale features is extracted using the transformer multiscaled head attention model. The proposed model employs different attention dimensions for each head of the transformer which enables it to provide attention at the multiscale level. Further, in addition to single modality, recognition performance using multiple modalities is examined. Extensive experiments demonstrate the superior performance of the proposed MsMHA-VTN with an overall accuracy of 88.22\% and 99.10\% on NVGesture and Briareo datasets, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。