统一模型实现手语翻译与字幕对齐,提升手语无障碍沟通效率。
Lost in Translation, Found in Embeddings: Sign Language Translation and Alignment
- 基于关键点与唇部图像的轻量视觉主干,保护用户隐私。
- 滑动感知器网络将视频特征映射为词级嵌入,精准对齐时序。
- 多任务联合训练实现跨语言泛化,支持英式与美式手语转换。
本文旨在构建一个统一模型,同时完成手语翻译(SLT)与手语-字幕对齐(SSA)任务。该模型可将连续手语视频转化为口语文本,并实现手语与字幕的精准时间对齐,对实际交流、大规模语料库建设及教育应用具有重要意义。方法包含三部分:(i) 轻量级视觉主干,从人体关键点和唇部区域图像中提取手动与非手动线索,保障签名人隐私;(ii) 滑动感知器映射网络,将连续视觉特征聚合为词级嵌入,弥合视觉与文本之间的鸿沟;(iii) 多任务可扩展训练策略,联合优化SLT与SSA,强化语言与时序对齐。模型在覆盖英国手语(BSL)和美国手语(ASL)的大规模手语-文本语料库(BOBSL与YouTube-SL-25)上进行预训练,实现在挑战性BOBSL(BSL)数据集上的最优性能,且在How2Sign(ASL)上展现出强零样本泛化能力与微调后翻译效果,证明了跨手语大规模翻译的可行性。
原文摘要 · Abstract (English)
Our aim is to develop a unified model for sign language understanding, that performs sign language translation (SLT) and sign-subtitle alignment (SSA). Together, these two tasks enable the conversion of continuous signing videos into spoken language text and also the temporal alignment of signing with subtitles -- both essential for practical communication, large-scale corpus construction, and educational applications. To achieve this, our approach is built upon three components: (i) a lightweight visual backbone that captures manual and non-manual cues from human keypoints and lip-region images while preserving signer privacy; (ii) a Sliding Perceiver mapping network that aggregates consecutive visual features into word-level embeddings to bridge the vision-text gap; and (iii) a multi-task scalable training strategy that jointly optimises SLT and SSA, reinforcing both linguistic and temporal alignment. To promote cross-linguistic generalisation, we pretrain our model on large-scale sign-text corpora covering British Sign Language (BSL) and American Sign Language (ASL) from the BOBSL and YouTube-SL-25 datasets. With this multilingual pretraining and strong model design, we achieve state-of-the-art results on the challenging BOBSL (BSL) dataset for both SLT and SSA. Our model also demonstrates robust zero-shot generalisation and finetuned SLT performance on How2Sign (ASL), highlighting the potential of scalable translation across different sign languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。