用骨架线索提升手语自监督表示,效果优于现有方法。
SignRep: Enhancing Self-Supervised Sign Representations
- 预训练时引入手语先验的骨骼线索,结合掩码自编码器。
- 在多个数据集上实现最优识别性能,且仅用单模态输入。
- 适合手语识别、字典检索与翻译任务,计算成本更低。
手语表征学习因手势复杂的时空特性及标注数据稀缺而面临挑战。现有方法或依赖通用视觉预训练模型(缺乏手语特异性),或采用复杂的多模态、多分支架构。为此,我们提出一种可扩展的自监督手语表征学习框架。在训练RGB模型时,引入重要的手语归纳先验,利用简单但关键的骨骼线索预训练掩码自编码器。这些手语特定先验、特征正则化与对抗性风格无关损失共同构建强大骨干网络。值得注意的是,模型推理阶段无需骨骼关键点,避免了基于关键点模型在下游任务中的局限性。微调后,在WLASL、ASL-Citizen和NMFs-CSL数据集上达到手语识别最优性能,采用更简洁的单模态架构。除识别外,冻结模型在手语字典检索和翻译任务中表现优异,超越标准MAE预训练与基于骨骼的表示;同时降低现有手语翻译模型的训练计算开销,在Phoenix2014T、CSL-Daily和How2Sign上保持强性能。
原文摘要 · Abstract (English)
Sign language representation learning presents unique challenges due to the complex spatio-temporal nature of signs and the scarcity of labeled datasets. Existing methods often rely either on models pre-trained on general visual tasks, that lack sign-specific features, or use complex multimodal and multi-branch architectures. To bridge this gap, we introduce a scalable, self-supervised framework for sign representation learning. We leverage important inductive (sign) priors during the training of our RGB model. To do this, we leverage simple but important cues based on skeletons while pretraining a masked autoencoder. These sign specific priors alongside feature regularization and an adversarial style agnostic loss provide a powerful backbone. Notably, our model does not require skeletal keypoints during inference, avoiding the limitations of keypoint-based models during downstream tasks. When finetuned, we achieve state-of-the-art performance for sign recognition on the WLASL, ASL-Citizen and NMFs-CSL datasets, using a simpler architecture and with only a single-modality. Beyond recognition, our frozen model excels in sign dictionary retrieval and sign translation, surpassing standard MAE pretraining and skeletal-based representations in retrieval. It also reduces computational costs for training existing sign translation models while maintaining strong performance on Phoenix2014T, CSL-Daily and How2Sign.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。