用分割掩码自监督学习,提升手语识别的细微动作捕捉能力
SignMAE: Segmentation-Driven Self-Supervised Learning for Sign Language Recognition

- 基于身体部位分割动态掩码,捕捉手部运动特征
- 在三个数据集上达到最优性能,用更少帧数和模态实现更高准确率
- 适合需要精准手语理解的场景,如无障碍交互系统
细微的手部差异使得手语识别极具挑战性,但现有方法多依赖通用动作数据集预训练的编码器,难以捕捉这些细粒度线索。本文提出一种面向手语识别的自监督预训练方法,通过基于分割的掩码机制,适应关键身体部位的存在与运动,而非将手部姿态视为静态视觉单元。该掩码重建目标显著提升了细粒度手语表征的学习能力。在WLASL、NMFs-CSL和Slovo数据集上,所提编码器均达到当前最优性能,在每实例和每类别Top-1准确率上均有提升,且所需输入帧数和模态数量少于同类方法。
原文摘要 · Abstract (English)
Subtle hand differences make sign language recognition challenging, yet many existing methods rely on encoders pretrained on generic action datasets that poorly capture such fine-grained cues. We propose a self-supervised pretraining method for sign language recognition that uses segmentation-based masking to adapt to the presence and motion of key body parts, rather than treating hand poses as static visual tokens. The resulting mask-and-reconstruct objective improves fine-grained sign representation learning. On WLASL, NMFs-CSL, and Slovo, our encoder achieves state-of-the-art performance, improving per-instance and per-class Top-1 accuracy while using fewer input frames and modalities than comparable encoders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。