arXiv:2509.14036cs.CLcs.AI2025-09

用对话提升手语翻译,无需复杂标注也能达到顶尖效果。

SSL-SSAW: Self-Supervised Learning with Sigmoid Self-Attention Weighting for Question-Based Sign Language Translation

  • 通过自监督学习融合对话与手语特征,动态加权关键信息。
  • 在两个新数据集上超越现有方法,对话辅助性能媲美甚至超过文字标注。
  • 适合关注真实场景手语翻译、降低标注成本的研究者。

手语翻译(SLT)有助于聋人与听力人士沟通,对话能提供重要上下文线索。本文提出基于问题的手语翻译(QB-SLT),利用自然发生的对话实现高效上下文整合。相比需人工标注的词元(gloss),对话更易获取且易于标注。核心挑战在于多模态特征对齐及利用问题上下文提升翻译质量。为此,我们提出交叉模态自监督学习结合Sigmoid自注意力加权(SSL-SSAW)融合方法:先用对比学习对齐多模态特征,再引入SSAW模块自适应提取问题与手语序列特征;同时利用自监督学习增强可得问题文本的表示能力。在新构建的CSL-Daily-QA与PHOENIX-2014T-QA数据集上评估,SSL-SSAW达到当前最优性能。值得注意的是,仅使用易获取的问题辅助即可实现或超越传统词元辅助的效果。可视化结果验证了对话引入对翻译质量的提升作用。

原文摘要 · Abstract (English)

Sign Language Translation (SLT) bridges the communication gap between deaf people and hearing people, where dialogue provides crucial contextual cues to aid in translation. Building on this foundational concept, this paper proposes Question-based Sign Language Translation (QB-SLT), a novel task that explores the efficient integration of dialogue. Unlike gloss (sign language transcription) annotations, dialogue naturally occurs in communication and is easier to annotate. The key challenge lies in aligning multimodality features while leveraging the context of the question to improve translation. To address this issue, we propose a cross-modality Self-supervised Learning with Sigmoid Self-attention Weighting (SSL-SSAW) fusion method for sign language translation. Specifically, we employ contrastive learning to align multimodality features in QB-SLT, then introduce a Sigmoid Self-attention Weighting (SSAW) module for adaptive feature extraction from question and sign language sequences. Additionally, we leverage available question text through self-supervised learning to enhance representation and translation capabilities. We evaluated our approach on newly constructed CSL-Daily-QA and PHOENIX-2014T-QA datasets, where SSL-SSAW achieved SOTA performance. Notably, easily accessible question assistance can achieve or even surpass the performance of gloss assistance. Furthermore, visualization results demonstrate the effectiveness of incorporating dialogue in improving translation quality.

手语翻译自监督学习对话建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。