arXiv:2603.16163cs.CVcs.CL2026-03

用统一注意力机制减少参数,提升手语识别效率。

STARK: Spatio-Temporal Attention for Representation of Keypoints for Continuous Sign Language Recognition

  • 设计时空联合注意力网络,同时建模关键点空间关系和时间动态。
  • 参数量减少70%-80%,在Phoenix-14T数据集上表现相当。
  • 适合追求高效轻量级手语识别系统的研究者与开发者。

连续手语识别(CSLR)是理解聋人语言的重要任务。现有基于关键点的方法通常采用时空编码:使用图卷积网络或注意力机制建模关键点间的空间交互,用一维卷积捕捉时间动态。但这类设计常导致编码器与解码器参数量巨大。本文提出一种统一的时空注意力网络,同时在空间(关键点间)和时间(局部窗口内)计算注意力分数,并聚合特征生成局部上下文感知的时空表示。所提编码器参数量比现有最先进模型减少约70%-80%,在Phoenix-14T数据集上性能相当。

原文摘要 · Abstract (English)

Continuous Sign Language Recognition (CSLR) is a crucial task for understanding the languages of deaf communities. Contemporary keypoint-based approaches typically rely on spatio-temporal encoding, where spatial interactions among keypoints are modeled using Graph Convolutional Networks or attention mechanisms, while temporal dynamics are captured using 1D convolutional networks. However, such designs often introduce a large number of parameters in both the encoder and the decoder. This paper introduces a unified spatio-temporal attention network that computes attention scores both spatially (across keypoints) and temporally (within local windows), and aggregates features to produce a local context-aware spatio-temporal representation. The proposed encoder contains approximately $70-80\%$ fewer parameters than existing state-of-the-art models while achieving comparable performance to keypoint-based methods on the Phoenix-14T dataset.

手语识别注意力机制轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。