arXiv:2505.10267cs.CVcs.LG2025-05被引 2

提出多模态手指拼写识别模型,提升手语命名识别准确率。

HandReader: Advanced Techniques for Efficient Fingerspelling Recognition

  • 设计三种新架构,融合视觉与关键点信息处理时序动作。
  • 在芝加哥和俄语数据集上达最新最好性能,准确率超现有方法。
  • 开源数据集与预训练模型,助力手语技术研究与应用。

手指拼写是手语中表示专有名词的重要组成部分,以快速手部动作为特征。尽管以往研究聚焦于视频时序维度处理,但识别精度仍有提升空间。本文提出HandReader,包含三种架构:HandReader$_{RGB}$采用新颖的时序自适应模块(TSAM)处理不同长度视频的RGB特征,保留关键时序信息;HandReader$_{KP}$基于提出的时序姿态编码器(TPE),对关键点张量进行2D/3D卷积,融合时空信息并累积坐标;此外还提出联合编码的HandReader_RGB+KP,结合视觉与关键点模态优势。各模型在ChicagoFSWild与ChicagoFSWild+数据集上达到当前最优,且在本文首次公开的俄语手指拼写数据集Znaki上表现优异。相关数据集与预训练模型已开源。

原文摘要 · Abstract (English)

Fingerspelling is a significant component of Sign Language (SL), allowing the interpretation of proper names, characterized by fast hand movements during signing. Although previous works on fingerspelling recognition have focused on processing the temporal dimension of videos, there remains room for improving the accuracy of these approaches. This paper introduces HandReader, a group of three architectures designed to address the fingerspelling recognition task. HandReader$_{RGB}$ employs the novel Temporal Shift-Adaptive Module (TSAM) to process RGB features from videos of varying lengths while preserving important sequential information. HandReader$_{KP}$ is built on the proposed Temporal Pose Encoder (TPE) operated on keypoints as tensors. Such keypoints composition in a batch allows the encoder to pass them through 2D and 3D convolution layers, utilizing temporal and spatial information and accumulating keypoints coordinates. We also introduce HandReader_RGB+KP - architecture with a joint encoder to benefit from RGB and keypoint modalities. Each HandReader model possesses distinct advantages and achieves state-of-the-art results on the ChicagoFSWild and ChicagoFSWild+ datasets. Moreover, the models demonstrate high performance on the first open dataset for Russian fingerspelling, Znaki, presented in this paper. The Znaki dataset and HandReader pre-trained models are publicly available.

手语识别手指拼写多模态关键点

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。