端到端学习手语视频的时空特征并翻译成文本。
Spatio-temporal Sign Language Representation and Translation
- 用统一模型同时学习手语的时空特征与翻译。
- 开发集上达到5±1的BLEU,测试集仅0.11±0.06。
- 适合研究跨模态表示与手语翻译的新方法。
本文介绍了DFKI-MLT团队在WMT-SLT 2022手语翻译任务中的提交,将瑞士德语手语(视频)翻译为德语(文本)。现有手语翻译系统多采用通用序列到序列架构,使用从视频帧中提取的特征作为输入,而非文本词嵌入。传统方法通常忽略时间特征。本文提出一种统一模型,联合学习时空特征表示与翻译,实现真正的端到端架构,预期对新数据集具有更好泛化能力。最佳系统在开发集上获得5±1 BLEU,但测试集性能降至0.11±0.06 BLEU。
原文摘要 · Abstract (English)
This paper describes the DFKI-MLT submission to the WMT-SLT 2022 sign language translation (SLT) task from Swiss German Sign Language (video) into German (text). State-of-the-art techniques for SLT use a generic seq2seq architecture with customized input embeddings. Instead of word embeddings as used in textual machine translation, SLT systems use features extracted from video frames. Standard approaches often do not benefit from temporal features. In our participation, we present a system that learns spatio-temporal feature representations and translation in a single model, resulting in a real end-to-end architecture expected to better generalize to new data sets. Our best system achieved $5\pm1$ BLEU points on the development set, but the performance on the test dropped to $0.11\pm0.06$ BLEU points.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。