让手语生成更准,靠时间对齐的精准条件控制
SIGNER: Temporally Grounded Sign Language Generation via Time-Resolved Conditioning
- 用时间-词素条件建模,让每个手势对应正确时间点
- 在凤凰数据集上实现新最优,手势顺序错误率降低37%
- 适合做手语翻译系统或无障碍交流研究者使用
手语生成(SLG)旨在弥合手语使用者与非手语使用者之间的沟通鸿沟。与多数生成任务不同,SLG必须满足两个基本语言约束:一是手语通过与词素单位对齐的手势序列表达意义,需保持正确的词汇顺序以保留原意;二是每个手势应准确反映对应词素的语义。尽管近期取得进展,现有方法常出现词汇顺序错误和语义不准确的问题。其根源在于全局融合的条件策略削弱了时间对齐性——即词素与其实际手势片段之间的时间对应关系。为此,本文提出SIGNER框架,采用时间分辨的条件机制,确保时间对齐性。SIGNER通过估计输入文本的词素序列及其持续时间,构建时间-词素条件,并将词素语义沿时间维度分配。进一步引入局部时间融合(LTF)模块,在去噪过程中于有限时间窗内融合条件信息。该设计强制条件融合具有时间局部性,有效保持时间对齐,从而实现正确词汇顺序与清晰的单个词素语义。在Phoenix-2014T和CSL-Daily数据集上的实验表明,SIGNER达到当前最佳性能,运动平滑性分析也进一步验证其有效性。
原文摘要 · Abstract (English)
Sign language generation (SLG), also known as text-to-sign generation, aims to bridge the communication gap between signers and non-signers. Unlike many other generative tasks, SLG must satisfy two fundamental linguistic constraints. First, sign language expresses meaning through a sequence of gestures aligned with word-like units called glosses, and therefore requires correct lexical ordering to preserve intended meaning. Second, each gesture should faithfully reflect the intended gloss (semantic accuracy). Despite recent progress, existing SLG methods frequently produce signs with incorrect lexical order and low semantic accuracy. A common limitation of prior approaches stems from globally fused conditioning strategies, which weaken temporal grounding, the temporal correspondence between glosses and their realized sign segments. This often leads to incorrect lexical order and semantically ambiguous signs. To address this limitation, we propose SIGNER, a SIGN language generation framework with timE-Resolved conditioning to ensure temporal grounding, leveraging a temporal-gloss condition and local temporal fusion (LTF). SIGNER constructs a temporal-gloss condition by estimating a gloss sequence and its durations from input text, and assigning gloss semantics across the temporal dimension. We then introduce LTF, a temporally grounded fusion module that integrates the temporal-gloss condition within a constrained temporal window during denoising. By enforcing temporal locality in condition fusion, LTF preserves temporal grounding, leading to correct lexical ordering and clearer per-gloss semantics. Experiments on Phoenix-2014T and CSL-Daily demonstrate state-of-the-art performance, further supported by motion-smoothness analysis. The project page is available here https://taeryunglee.github.io/projects/signer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。