统一手语翻译与生成,实现文本与手语互转。
Bridging the Gap Between Semantics and Reconstruction:Unifying Sign Language Translation and Production

- 用共享手语分词器同时捕捉语义和动作细节。
- 单一模型可双向转换:手语转文本或文本转手语。
- 在生成精度上优于现有方法,适合跨模态研究者。
近期手语研究趋向于将孤立手语识别(ISLR)、连续手语识别(CSLR)与手语翻译(SLT)等任务统一在一个框架中,取得显著进展。与此同时,从文本生成手语序列的手语生成(SLP)也日益受到关注。这引出关键问题:能否在单一框架内统一手语理解与生成?相较于统一理解子任务,该问题更具挑战性——因为SLT与SLP的映射方向相反。为此,本文提出Uni-SLTP,包含两个核心组件:(1) 共享手语分词器,将手语序列转化为离散符号与潜在表示,兼顾语义抽象与动作重建;(2) 统一自回归生成模型,将两类任务均建模为条件序列生成。在多个公开数据集上的实验表明,Uni-SLTP在手语生成任务中实现了更优的动作准确性,同时保持了竞争力的翻译性能。
原文摘要 · Abstract (English)
Recent advances in sign language (SL) research have shown a trend toward unifying multiple sign language understanding (SLU) subtasks, such as isolated sign language recognition (ISLR), continuous sign language recognition (CSLR), and sign language translation (SLT), within a single framework, leading to substantial progress. Meanwhile, sign language production (SLP), which generates sign sequences from text, has also attracted growing attention. This naturally raises an important question: can sign language understanding and production be unified within a single framework? Compared with unifying SLU subtasks, this problem is substantially more challenging. Existing SLU tasks largely share the same direction of mapping, namely from sign inputs to linguistic outputs, whereas SLT and SLP lie in opposite directions of sign-text mapping. A unified framework must therefore address two key challenges: (1) bridging the modality gap between continuous sign motions and discrete text tokens through a shared sign tokenizer that supports both linguistic abstraction and motion reconstruction; and (2) learning a single conditional autoregressive model that can take either sign or text as input and generate the corresponding target sequence in the opposite modality. To this end, we propose Uni-SLTP, a unified framework for SLT and SLP with two key components: (1) a shared sign tokenizer that converts sign sequences into discrete tokens and latent representations, capturing both semantic and reconstructive information; and (2) a unified autoregressive generation model that formulates both tasks as conditional sequence generation. Experiments on widely used public datasets show that Uni-SLTP achieves superior motion accuracy for SLP while maintaining competitive SLT performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。