用上下文提升手语翻译准确率,效果优于现有方法。
Lost in Translation, Found in Context: Sign Language Translation with Contextual Cues
- 融合视频、字幕、前句翻译和伪词注释等上下文信息。
- 在BOBSL数据集上显著提升翻译质量,超越已有成果。
- 方法通用性强,适用于英式与美式手语数据集。
本研究旨在将连续手语翻译为口语文字。受人类译员依赖上下文的启发,我们提出一种新框架,将视觉手语识别特征与额外上下文线索结合:(i) 背景节目的字幕描述,(ii) 前一句的翻译结果,以及 (iii) 手语的伪词注释。这些信息自动提取后,与视觉特征一同输入预训练大语言模型(LLM),并进行微调以生成文本形式的口语翻译。通过大量消融实验,验证了每种输入线索对翻译性能的积极贡献。我们在当前最大的英式手语数据集BOBSL上训练并评估该方法,结果显示其翻译质量显著优于此前报告的结果及多个基准的最先进方法。此外,我们将该方法扩展至美式手语数据集How2Sign,同样取得具有竞争力的效果。
原文摘要 · Abstract (English)
Our objective is to translate continuous sign language into spoken language text. Inspired by the way human interpreters rely on context for accurate translation, we incorporate additional contextual cues together with the signing video, into a new translation framework. Specifically, besides visual sign recognition features that encode the input video, we integrate complementary textual information from (i) captions describing the background show, (ii) translation of previous sentences, as well as (iii) pseudo-glosses transcribing the signing. These are automatically extracted and inputted along with the visual features to a pre-trained large language model (LLM), which we fine-tune to generate spoken language translations in text form. Through extensive ablation studies, we show the positive contribution of each input cue to the translation performance. We train and evaluate our approach on BOBSL -- the largest British Sign Language dataset currently available. We show that our contextual approach significantly enhances the quality of the translations compared to previously reported results on BOBSL, and also to state-of-the-art methods that we implement as baselines. Furthermore, we demonstrate the generality of our approach by applying it also to How2Sign, an American Sign Language dataset, and achieve competitive results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。