arXiv:2503.02360cs.CVcs.AI2025-03被引 10

构建首个大规模孟加拉手语数据集,用新编码方法提升识别准确率。

BdSLW401: Transformer-Based Word-Level Bangla Sign Language Recognition Using Relative Quantization Encoding (RQE)

  • 用生理参考点量化手势轨迹,减少空间变化干扰注意力
  • 在多个数据集上降低44.3%的词错误率,显著提升性能
  • 适合研究低资源语言手语识别与可解释性模型的学者

针对孟加拉手语(BdSL)这类低资源语言的手语识别(SLR)面临签者差异、视角变化和标注数据稀缺问题。本文提出BdSLW401,一个大规模多视角词级手语数据集,包含401个手语词、102,176段视频样本,来自18位签者在正面与侧面视角下的录制。为提升基于Transformer的SLR效果,提出相对量化编码(RQE),通过将关键点锚定于生理参考点并量化运动轨迹,减少空间变异,优化注意力分配。RQE在WLASL100上实现44.3%的词错误率(WER)下降,在SignBD-200上下降21.0%,并在BdSLW60和SignBD-90中取得显著提升。然而,固定量化在大规模数据集(如WLASL2000)上表现不足,表明需自适应编码策略。进一步提出的RQE-SF变体通过稳定肩部关键点提升姿态一致性,但小幅牺牲侧视识别效果。注意力图分析显示,RQE使模型更聚焦于主要构音特征(手指、手腕)及更具区分性的帧,而非整体姿势变化。本工作引入了BdSLW401,并验证了RQE增强结构化嵌入的有效性,推动了低资源语言下Transformer手语识别的发展,为后续研究设定了基准。

原文摘要 · Abstract (English)

Sign language recognition (SLR) for low-resource languages like Bangla suffers from signer variability, viewpoint variations, and limited annotated datasets. In this paper, we present BdSLW401, a large-scale, multi-view, word-level Bangla Sign Language (BdSL) dataset with 401 signs and 102,176 video samples from 18 signers in front and lateral views. To improve transformer-based SLR, we introduce Relative Quantization Encoding (RQE), a structured embedding approach anchoring landmarks to physiological reference points and quantize motion trajectories. RQE improves attention allocation by decreasing spatial variability, resulting in 44.3% WER reduction in WLASL100, 21.0% in SignBD-200, and significant gains in BdSLW60 and SignBD-90. However, fixed quantization becomes insufficient on large-scale datasets (e.g., WLASL2000), indicating the need for adaptive encoding strategies. Further, RQE-SF, an extended variant that stabilizes shoulder landmarks, achieves improvements in pose consistency at the cost of small trade-offs in lateral view recognition. The attention graphs prove that RQE improves model interpretability by focusing on the major articulatory features (fingers, wrists) and the more distinctive frames instead of global pose changes. Introducing BdSLW401 and demonstrating the effectiveness of RQE-enhanced structured embeddings, this work advances transformer-based SLR for low-resource languages and sets a benchmark for future research in this area.

手语识别低资源语言注意力机制数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。