arXiv:2506.13156cs.CV2025-06被引 1

用图模型生成手语连贯过渡视频,让手势更自然流畅。

StgcDiff: Spatial-Temporal Graph Condition Diffusion for Sign Language Transition Generation

  • 基于时空图结构建模手语动作,捕捉复杂空间时间关系。
  • 在三个数据集上生成视频质量优于现有方法,过渡更平滑。
  • 适合手语合成、人机交互等需要自然手势生成的场景。

手语过渡生成旨在将离散的手语片段转换为连续的手语视频,通过合成平滑的过渡来提升视觉连贯性和语义准确性。然而,大多数现有方法仅简单拼接孤立手势,导致生成视频视觉不连贯且语义不准确。与文本语言不同,手语本身蕴含丰富的时空信息,建模难度更高。为此,我们提出StgcDiff,一种基于图的条件扩散框架,通过捕捉手语独特的时空依赖关系,生成两个离散手势间的自然过渡。首先,训练一个编码器-解码器架构,学习时空骨骼序列的结构感知表示;其次,优化一个以预训练编码器表示为条件的扩散去噪器,从噪声中预测过渡帧。此外,设计了Sign-GCN模块作为核心组件,有效建模时空特征。在PHOENIX14T、USTC-CSL100和USTC-SLR500三个数据集上的大量实验表明,该方法性能显著优于现有方法。

原文摘要 · Abstract (English)

Sign language transition generation seeks to convert discrete sign language segments into continuous sign videos by synthesizing smooth transitions. However,most existing methods merely concatenate isolated signs, resulting in poor visual coherence and semantic accuracy in the generated videos. Unlike textual languages,sign language is inherently rich in spatial-temporal cues, making it more complex to model. To address this,we propose StgcDiff, a graph-based conditional diffusion framework that generates smooth transitions between discrete signs by capturing the unique spatial-temporal dependencies of sign language. Specifically, we first train an encoder-decoder architecture to learn a structure-aware representation of spatial-temporal skeleton sequences. Next, we optimize a diffusion denoiser conditioned on the representations learned by the pre-trained encoder, which is tasked with predicting transition frames from noise. Additionally, we design the Sign-GCN module as the key component in our framework, which effectively models the spatial-temporal features. Extensive experiments conducted on the PHOENIX14T, USTC-CSL100,and USTC-SLR500 datasets demonstrate the superior performance of our method.

手语生成扩散模型图神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。