用多任务变换器实现手语到文本的精准翻译,提升聋人沟通效率。
A multitask transformer to sign language translation using motion gesture primitives
- 引入手势基元与密集运动表示,融合运动学信息增强手语特征
- 在CoL-SLTD数据集上达到72.64% BLEU-4(Split 1),显著超越现有方法
- 适用于手语翻译、无障碍通信系统开发,尤其适合语音-手语交互研究者
聋人群体因缺乏有效沟通方式而面临主要社会隔阂。手语作为其主要交流工具,缺乏书面形式,导致当前核心挑战在于时空手语表示与自然语言之间的自动翻译。现有方法多采用编码器-解码器架构,依赖注意力模块捕捉非线性对应关系,但常需复杂训练与结构,且受限于视频序列中的冗余背景信息。本文提出一种多任务变压器架构,引入词素学习表示以实现更优翻译。该方法包含密集运动表示,强化手势特征并融入运动学信息,有助于抑制背景干扰并利用手语几何特性;同时结合时空表示,促进手势与词素间的对齐。在CoL-SLTD数据集上,该方法在Split 1达到72.64% BLEU-4,Split 2为14.64%;在RWTH-PHOENIX-Weather 2014 T数据集上也取得11.58%的竞争力BLEU-4得分。
原文摘要 · Abstract (English)
The absence of effective communication the deaf population represents the main social gap in this community. Furthermore, the sign language, main deaf communication tool, is unlettered, i.e., there is no formal written representation. In consequence, main challenge today is the automatic translation among spatiotemporal sign representation and natural text language. Recent approaches are based on encoder-decoder architectures, where the most relevant strategies integrate attention modules to enhance non-linear correspondences, besides, many of these approximations require complex training and architectural schemes to achieve reasonable predictions, because of the absence of intermediate text projections. However, they are still limited by the redundant background information of the video sequences. This work introduces a multitask transformer architecture that includes a gloss learning representation to achieve a more suitable translation. The proposed approach also includes a dense motion representation that enhances gestures and includes kinematic information, a key component in sign language. From this representation it is possible to avoid background information and exploit the geometry of the signs, in addition, it includes spatiotemporal representations that facilitate the alignment between gestures and glosses as an intermediate textual representation. The proposed approach outperforms the state-of-the-art evaluated on the CoL-SLTD dataset, achieving a BLEU-4 of 72,64% in split 1, and a BLEU-4 of 14,64% in split 2. Additionally, the strategy was validated on the RWTH-PHOENIX-Weather 2014 T dataset, achieving a competitive BLEU-4 of 11,58%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。