用目标语句嵌入代替人工标注词元,实现无需词元的聋哑人手语翻译。
Sign Language Translation with Sentence Embedding Supervision
- 训练时用目标句子的嵌入作为监督信号替代词元标注。
- 在德国和美国手语数据集上性能超越现有无词元方法。
- 支持多语言,适合缺乏标注数据的手语翻译研究者。
当前最先进的手语翻译系统依赖词元标注来辅助学习,但这类标注数据稀缺且跨数据集差异大。本文提出一种新方法:在训练时使用目标句子的嵌入作为监督信号,取代词元标注。该监督信号来自原始文本数据,无需人工标注。由于方法天然支持多语言,我们在涵盖德语(PHOENIX-2014T)和美式手语(How2Sign)的数据集上进行了评估,并尝试了单语和多语言嵌入及翻译系统。结果表明,该方法显著优于其他无词元方法,在无额外预训练数据的情况下达到新的最优水平,大幅缩小了无词元与依赖词元系统间的差距。
原文摘要 · Abstract (English)
State-of-the-art sign language translation (SLT) systems facilitate the learning process through gloss annotations, either in an end2end manner or by involving an intermediate step. Unfortunately, gloss labelled sign language data is usually not available at scale and, when available, gloss annotations widely differ from dataset to dataset. We present a novel approach using sentence embeddings of the target sentences at training time that take the role of glosses. The new kind of supervision does not need any manual annotation but it is learned on raw textual data. As our approach easily facilitates multilinguality, we evaluate it on datasets covering German (PHOENIX-2014T) and American (How2Sign) sign languages and experiment with mono- and multilingual sentence embeddings and translation systems. Our approach significantly outperforms other gloss-free approaches, setting the new state-of-the-art for data sets where glosses are not available and when no additional SLT datasets are used for pretraining, diminishing the gap between gloss-free and gloss-dependent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。