用时空变压器提升手语翻译,兼顾局部动作和长距离语义
Spatio-temporal transformer to support automatic sign language translation
- 设计多卷积与注意力融合的时空变换器,捕捉手势时空特征
- 在CoL-SLTD上达46.84% BLEU4,PHOENIX14T上达30.77%
- 适合需要高精度手语翻译的无障碍应用开发者
手语翻译(SLT)系统通过建立手语与口语之间的对应关系,帮助听障人士沟通。但该任务因手语变体多样、语言结构复杂且表达丰富而极具挑战。现有计算方法虽有进展,但仍难以覆盖手势差异并处理长序列翻译。本文提出一种基于Transformer的架构,通过多卷积与注意力机制联合编码时空运动手势,有效保留局部与长程空间信息。在哥伦比亚手语翻译数据集(CoL-SLTD)上,性能优于基线模型,取得46.84%的BLEU4得分;在RWTH-PHOENIX-Weather-2014T(PHOENIX14T)上达到30.77%的BLEU4,验证了其对真实世界变化的鲁棒性与有效性。
原文摘要 · Abstract (English)
Sign Language Translation (SLT) systems support hearing-impaired people communication by finding equivalences between signed and spoken languages. This task is however challenging due to multiple sign variations, complexity in language and inherent richness of expressions. Computational approaches have evidenced capabilities to support SLT. Nonetheless, these approaches remain limited to cover gestures variability and support long sequence translations. This paper introduces a Transformer-based architecture that encodes spatio-temporal motion gestures, preserving both local and long-range spatial information through the use of multiple convolutional and attention mechanisms. The proposed approach was validated on the Colombian Sign Language Translation Dataset (CoL-SLTD) outperforming baseline approaches, and achieving a BLEU4 of 46.84%. Additionally, the proposed approach was validated on the RWTH-PHOENIX-Weather-2014T (PHOENIX14T), achieving a BLEU4 score of 30.77%, demonstrating its robustness and effectiveness in handling real-world variations
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。