arXiv:2604.10413cs.SD2026-04中稿 · ICPR 2026

让手语的语气直接融入语音,实现更自然的跨模态表达。

Sign-to-Speech Prosody Transfer via Sign Reconstruction-based GAN

  • 用对抗学习和重建损失,无需成对标注即可训练
  • 能准确传递手语的情感语气,合成语音更贴近原意
  • 适合手语语音转换、无障碍沟通系统开发者

深度学习提升了手语到文本的翻译能力,使非手语者更容易理解手语内容。若目标是语音交流,传统方法先将手语转为文本再通过文本转语音(TTS)合成语音,但这一两阶段流程会丢失手语中丰富的非语言信息。为此,本文提出新任务:手语到语音的韵律迁移,旨在捕捉手语中的全局韵律特征,并直接注入合成语音中。主要挑战在于手语与语音对齐需专家知识,标注成本极高,难以构建大规模平行数据集。为此,我们提出SignRecGAN框架,利用单模态数据集,通过对抗学习和重建损失实现可扩展训练。同时提出S2PFormer模型架构,在保留现有TTS表达力的基础上,支持注入手语来源的韵律信息。大量实验表明,该方法能生成忠实反映手语情感内容的语音,为更自然的手语交流开辟新路径。代码将在录用后公开。

原文摘要 · Abstract (English)

Deep learning models have improved sign language-to-text translation and made it easier for non-signers to understand signed messages. When the goal is spoken communication, a naive approach is to convert signed messages into text and then synthesize speech via Text-to-Speech (TTS). However, this two-stage pipeline inevitably treat text as a bottleneck representation, causing the loss of rich non-verbal information originally conveyed in the signing. To address this limitation, we propose a novel task, \emph{Sign-to-Speech Prosody Transfer}, which aims to capture the global prosodic nuances expressed in sign language and directly integrate them into synthesized speech. A major challenge is that aligning sign and speech requires expert knowledge, making annotation extremely costly and preventing the construction of large parallel corpora. To overcome this, we introduce \emph{SignRecGAN}, a scalable training framework that leverages unimodal datasets without cross-modal annotations through adversarial learning and reconstruction losses. Furthermore, we propose \emph{S2PFormer}, a new model architecture that preserves the expressive power of existing TTS models while enabling the injection of sign-derived prosody into the synthesized speech. Extensive experiments demonstrate that the proposed method can synthesize speech that faithfully reflects the emotional content of sign language, thereby opening new possibilities for more natural sign language communication. Our code will be available upon acceptance.

手语转换语音合成韵律迁移多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。