arXiv:2606.21234cs.CV2026-06

分词生成手语,让动作更自然连贯。

Context-Aware Autoregressive Diffusion for Gloss-Wise Sign Language Production

论文配图:Context-Aware Autoregressive Diffusion for Gloss-Wise Sign Language Production
图 1 · 摘自论文原文
  • 按词单位逐步生成手语,结合语义与动作上下文。
  • 在Phoenix-T和CSL-Daily上准确率提升,动作更流畅。
  • 适合需要精细控制手语动作的研究与应用。

为生成自然准确的句级手语,合成基本语义单元“gloss”至关重要。现有手语生成(SLP)方法多一次性生成整串序列,虽高效但易产生时间漂移和手势模糊,且难以精确控制单个gloss。本文提出上下文感知的分词自回归扩散模型GARD,通过同时依赖语义(语言)和运动(动作)上下文来建模协同发音。为确保gloss间动作连续性,GARD引入两项策略:i)跨词过渡引导,基于梯度对齐词间边界动作,保证姿态一致;ii)全局动作调谐器,根据调整后的边界姿态优化整个动作序列。在Phoenix-T和CSL-Daily数据集上的实验表明,GARD在语言准确性和动作相似性上均优于现有方法。

原文摘要 · Abstract (English)

To generate natural and accurate sentence-level sign language, synthesizing the "gloss", the fundamental semantic unit, is essential. However, most current sign-language production (SLP) methods generate entire sequences at once. While this end-to-end approach is often efficient, it is prone to temporal drift and hand motion blur as sentences get longer, and fails to accurately control individual glosses. In this paper, we propose the Context-aware Gloss-wise AutoRegressive Diffusion model (GARD), a gloss-wise diffusion framework that models coarticulation by conditioning on both semantic (linguistic) and kinematic (motion) contexts. To ensure natural continuity between gloss motions, GARD introduces two additional strategies: i) Inter-Gloss Transition Guidance, which applies gradient-based guidance to kinematically align inter-gloss boundaries and ensure seamless pose consistency. ii) Global Motion Harmonizer, refining the entire gloss motion sequence based on the boundary poses adjusted by Inter-Gloss Transition Guidance. Extensive experiments on Phoenix-T and CSL-Daily datasets demonstrate that GARD achieves superior performance over existing SLP methods in terms of both linguistic accuracy and motion similarity.

手语生成扩散模型自回归动作连续

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。