arXiv:2505.24355cs.CL2025-05ACL被引 11

无需词典的多语言手语翻译模型,支持10种手语互译。

Multilingual Gloss-free Sign Language Translation: Towards Building a Sign Language Foundation Model

  • 设计双CTC目标,直接从视频中识别手语动作并生成文本。
  • 在3个基准上实现与顶尖方法相当的性能,覆盖10种手语。
  • 适用于多对一、多对多手语翻译,推动无障碍沟通发展。

手语翻译(SLT)旨在将手语视频转换为口语文字,弥合手语使用者与口语者之间的沟通鸿沟。现有研究多聚焦单一手语到单一语言的一对一翻译,而利用多语言资源可缓解低资源问题并提升可及性。然而,由于手语与口语间存在语言冲突和对齐难题,多语言手语翻译(MLSLT)仍处于空白状态。为此,我们提出一种无词典的多语言模型,采用双重CTC目标实现逐标记的手语识别与口语文本生成。该模型支持10种手语,可处理一对一、多对一及多对多的翻译任务,在三个主流基准——多语言SP-10、PHOENIX14T和CSL-Daily上达到与当前最优方法相当的性能。

原文摘要 · Abstract (English)

Sign Language Translation (SLT) aims to convert sign language (SL) videos into spoken language text, thereby bridging the communication gap between the sign and the spoken community. While most existing works focus on translating a single sign language into a single spoken language (one-to-one SLT), leveraging multilingual resources could mitigate low-resource issues and enhance accessibility. However, multilingual SLT (MLSLT) remains unexplored due to language conflicts and alignment difficulties across SLs and spoken languages. To address these challenges, we propose a multilingual gloss-free model with dual CTC objectives for token-level SL identification and spoken text generation. Our model supports 10 SLs and handles one-to-one, many-to-one, and many-to-many SLT tasks, achieving competitive performance compared to state-of-the-art methods on three widely adopted benchmarks: multilingual SP-10, PHOENIX14T, and CSL-Daily.

手语翻译多语言无词典视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。