arXiv:2411.16765cs.CLcs.CV2024-11ACL被引 25

用自监督学习构建手语上下文表征,提升多任务性能。

SHuBERT: Self-Supervised Sign Language Representation Learning via Multi-Stream Cluster Prediction

  • 通过多流聚类预测实现手语视频的自监督预训练
  • 在1000小时美式手语数据上达到多任务最优效果
  • 适合手语识别、翻译与拼写检测等研究者参考

手语处理长期依赖特定任务模型,限制了跨任务迁移。现有预训练方法或采用有监督方式,无法利用无标签数据;或仅关注单帧/片段表示,忽略时序关系。我们提出SHuBERT(Sign Hidden-Unit BERT),基于约1000小时美式手语视频,通过多流聚类预测目标,适配多模态视觉输入,学习手部、面部与身体姿态流的联合表征。该模型在手语翻译、孤立手语识别与手指拼写检测等多个任务上均取得当前最优表现。

原文摘要 · Abstract (English)

Sign language processing has traditionally relied on task-specific models, limiting the potential for transfer learning across tasks. Pre-training methods for sign language have typically focused on either supervised pre-training, which cannot take advantage of unlabeled data, or context-independent (frame or video segment) representations, which ignore the effects of relationships across time in sign language. We introduce SHuBERT (Sign Hidden-Unit BERT), a self-supervised contextual representation model learned from approximately 1,000 hours of American Sign Language video. SHuBERT adapts masked token prediction objectives to multi-stream visual sign language input, learning to predict multiple targets corresponding to clustered hand, face, and body pose streams. SHuBERT achieves state-of-the-art performance across multiple tasks including sign language translation, isolated sign language recognition, and fingerspelling detection.

手语识别自监督学习多模态表征BERT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。