arXiv:2411.17248cs.CV2024-11被引 6

用扩散模型生成更多样化的手语翻译,提升准确性与自然度。

DiffSLT: Enhancing Diversity in Sign Language Translation via Diffusion Model

  • 采用扩散模型从噪声生成目标文本表示,结合视频特征实现多样输出。
  • 在两个数据集上多样性显著提升,翻译质量达当前最优水平。
  • 适合关注手语翻译多样性与生成质量的研究者与开发者。

手语翻译(SLT)面临将手语视频转化为自然语言的挑战。以往研究侧重准确率而忽视多样性,但多样性对处理词汇与句法歧义至关重要,可能同样有益于手语翻译。本文提出DiffSLT,一种全新的无词元(gloss-free)SLT框架,利用扩散模型在保留手语语义的前提下生成多样化翻译结果。DiffSLT通过条件化输入视频的视觉特征,将随机噪声逐步转化为目标潜在表示。为增强视觉条件引导,设计了引导融合模块(Guidance Fusion Module),充分提取多层级时空视觉信息。此外,提出DiffSLT-P变体,同时使用伪词元和视觉特征进行条件控制,提供关键文本引导并缩小模态差距。实验表明,DiffSLT与DiffSLT-P在两个主流SLT数据集上均显著提升多样性,并达到当前最佳性能,显著改善翻译质量。

原文摘要 · Abstract (English)

Sign language translation (SLT) is challenging, as it involves converting sign language videos into natural language. Previous studies have prioritized accuracy over diversity. However, diversity is crucial for handling lexical and syntactic ambiguities in machine translation, suggesting it could similarly benefit SLT. In this work, we propose DiffSLT, a novel gloss-free SLT framework that leverages a diffusion model, enabling diverse translations while preserving sign language semantics. DiffSLT transforms random noise into the target latent representation, conditioned on the visual features of input video. To enhance visual conditioning, we design Guidance Fusion Module, which fully utilizes the multi-level spatiotemporal information of the visual features. We also introduce DiffSLT-P, a DiffSLT variant that conditions on pseudo-glosses and visual features, providing key textual guidance and reducing the modality gap. As a result, DiffSLT and DiffSLT-P significantly improve diversity over previous gloss-free SLT methods and achieve state-of-the-art performance on two SLT datasets, thereby markedly improving translation quality.

手语翻译扩散模型多样性生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。