用变换器模型实现手语连贯序列的精准分段。
Hands-On: Segmenting Individual Signs from Continuous Sequences
- 将分段任务转为BIO标签序列建模,捕捉手语时间动态。
- 在DGS语料库上达到当前最佳性能,优于已有方法。
- 结合手部特征与3D角度信息,适合手语识别研究者。
本研究解决连续手语分割这一关键挑战,对手语翻译与数据标注具有重要意义。提出一种基于变换器的架构,建模手语的时间动态性,并将分割问题转化为使用开始-中间-结束(BIO)标记方案的序列标注任务。方法融合了HaMeR手部特征及3D角度信息。大量实验表明,该模型在DGS语料库上达到当前最优表现,其特征在BSLCorpus上亦超越先前基准。
原文摘要 · Abstract (English)
This work tackles the challenge of continuous sign language segmentation, a key task with huge implications for sign language translation and data annotation. We propose a transformer-based architecture that models the temporal dynamics of signing and frames segmentation as a sequence labeling problem using the Begin-In-Out (BIO) tagging scheme. Our method leverages the HaMeR hand features, and is complemented with 3D Angles. Extensive experiments show that our model achieves state-of-the-art results on the DGS Corpus, while our features surpass prior benchmarks on BSLCorpus.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。