arXiv:2510.19398cs.CL2025-10被引 4

用跨语言多模态嵌入实现无需文本的多语种手语翻译

SONAR-SLT: Multilingual Sign Language Translation via Language-Agnostic Sentence Embedding Supervision

  • 用多语言文本语音训练的跨语言嵌入监督手语翻译
  • 低资源场景下比纯文本嵌入提升显著,BLEURT得分更高
  • 适合想做多语种手语系统或数据少的研究者

手语翻译通常依赖单一口语文本进行训练,限制了可扩展性和跨语言泛化能力。先前方法虽用文本句嵌入替代词位监督,但仍局限于特定语言和模态。本文采用在多种语言的文本与语音上训练的跨语言、多模态嵌入来监督手语翻译,实现直接的多语种翻译。为应对数据稀缺问题,提出一种耦合增强方法,结合多语言目标增强(即多种语言翻译)与视频级扰动,提升模型鲁棒性。实验表明,在所有设置下均优于仅使用文本句嵌入的监督方式,尤其在低资源场景下提升更明显。结果证明,结合跨语言嵌入监督与耦合增强,能提供一种可扩展且语义稳健的替代传统训练方案。

原文摘要 · Abstract (English)

Sign language translation (SLT) is typically trained with text in a single spoken language, which limits scalability and cross-language generalization. Earlier approaches have replaced gloss supervision with text-based sentence embeddings, but up to now, these remain tied to a specific language and modality. In contrast, here we employ language-agnostic, multimodal embeddings trained on text and speech from multiple languages to supervise SLT, enabling direct multilingual translation. To address data scarcity, we propose a coupled augmentation method that combines multilingual target augmentations (i.e. translations into many languages) with video-level perturbations, improving model robustness. Experiments show consistent BLEURT gains over text-only sentence embedding supervision, with larger improvements in low-resource settings. Our results demonstrate that language-agnostic embedding supervision, combined with coupled augmentation, provides a scalable and semantically robust alternative to traditional SLT training.

手语翻译多模态跨语言增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。