arXiv:2509.10845cs.CLcs.MM2025-09被引 1

无需词素标注,直接从文本生成手语动作序列

Text2Sign Diffusion: A Generative Approach for Gloss-Free Sign Language Production

  • 采用扩散模型联合噪声符号码与文本生成手语帧
  • 在PHOENIX14T和How2Sign上达到当前最佳性能
  • 适合无词素标注场景,提升手语生成泛化能力

手语生成(SLP)旨在将口语句子转换为手语的动作序列,弥合沟通鸿沟,促进聋哑人群的数字包容性。现有方法通常依赖词素(gloss),即手语词汇或短语的符号表示,作为中间步骤,这限制了SLP的灵活性与泛化能力,因为词素标注常不可用且具语言特异性。为此,我们提出一种全新的基于扩散的生成方法——Text2Sign Diffusion(Text2SignDiff),实现无词素的SLP。具体而言,提出一种无词素的潜在扩散模型,联合噪声潜在手语码与口语文本进行生成,通过非自回归的迭代去噪过程减少误差累积。同时设计跨模态手语对齐器,学习视觉与文本内容的共享潜在空间,支持条件扩散过程,实现更准确、上下文相关的手语生成。在常用数据集PHOENIX14T和How2Sign上的大量实验表明,该方法有效性显著,达到当前最优性能。

原文摘要 · Abstract (English)

Sign language production (SLP) aims to translate spoken language sentences into a sequence of pose frames in a sign language, bridging the communication gap and promoting digital inclusion for deaf and hard-of-hearing communities. Existing methods typically rely on gloss, a symbolic representation of sign language words or phrases that serves as an intermediate step in SLP. This limits the flexibility and generalization of SLP, as gloss annotations are often unavailable and language-specific. Therefore, we present a novel diffusion-based generative approach - Text2Sign Diffusion (Text2SignDiff) for gloss-free SLP. Specifically, a gloss-free latent diffusion model is proposed to generate sign language sequences from noisy latent sign codes and spoken text jointly, reducing the potential error accumulation through a non-autoregressive iterative denoising process. We also design a cross-modal signing aligner that learns a shared latent space to bridge visual and textual content in sign and spoken languages. This alignment supports the conditioned diffusion-based process, enabling more accurate and contextually relevant sign language generation without gloss. Extensive experiments on the commonly used PHOENIX14T and How2Sign datasets demonstrate the effectiveness of our method, achieving the state-of-the-art performance.

手语生成扩散模型无词素

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。