arXiv:2606.15889cs.CV2026-06

让语音同步手势既语义精准又保留个人风格。

SiGnature: Explicit Motion Diffusion for Stylized Semantic Gesture

论文配图:SiGnature: Explicit Motion Diffusion for Stylized Semantic Gesture
图 1 · 摘自论文原文
  • 在显式运动空间中设计动作融合机制,直接注入任意手势序列。
  • 无需重训练即可生成复杂语义手势,且保持自然流畅的肢体动态。
  • 适合需要个性化语音手势生成的研究与应用,如虚拟主播、人机交互。

尽管近期语音同步手势生成在节奏同步方面取得显著进展,但如何生成既具语义意义又忠实于说话者独特非语言风格的手势仍是开放挑战。语义手势(如象形动作或指指点点)在数据中统计稀疏,难以在标准生成模型中有效学习。我们提出 SiGnature,一种面向风格化与语义手势生成的框架,实现了精确语义控制与高保真风格保留的统一。不同于依赖隐式嵌入表示的主流方法,SiGnature 在显式的联合旋转空间中运行。其核心贡献是无需训练的推理机制——联合运动融合(JMI),可直接将任意外部运动序列(尤其是真实场景中的语义手势)注入扩散过程。JMI 自动识别传递语义动作的特定活动关节,并将其注入生成过程,其余身体动态(包括姿态与运动流)则由扩散主干网络根据目标说话者的预学风格合成。该方法支持任意动作的即插即用集成,无需重训练,避免了剪切拼接类方法常见的‘怪异融合’伪影。大量实验与感知评估表明,SiGnature 在语义动作控制方面表现更优,同时保持自然流畅的语音同步手势生成,并完整保留说话者个体特征,优于当前最先进基线。

原文摘要 · Abstract (English)

While recent advances in co-speech gesture generation have achieved impressive rhythmic synchronization, synthesizing gestures that are both semantically meaningful and faithful to a speaker's unique non-verbal style remains an open challenge. Semantic gestures, such as iconic shapes or deictic pointing, are statistically sparse, making them difficult to learn effectively within standard generative models. We present SiGnature, a framework for Stylized and Semantic Gesture generation that reconciles precise semantic control with high-fidelity style preservation. Unlike prevalent methods that rely on entangled latent representations, SiGnature operates in an explicit joint-rotation space. This design enables our core contribution, Joint Motion Integration (JMI), a training-free inference mechanism capable of injecting any external motion sequence, particularly in-the-wild semantic gestures, directly into the diffusion process. JMI automatically identifies the specific ``active joints'' conveying a semantic action and injects them into the generation, while relying on the diffusion backbone to synthesize the remaining body dynamics, including posture and flow, in accordance with the pre-learned style of the target speaker. This allows for the plug-and-play integration of arbitrary motions, including complex semantic gestures, without retraining or introducing the ``Frankenstein'' artifacts typical of cut-and-paste methods. Extensive experiments and perceptual studies demonstrate that SiGnature offers superior semantic motion control while maintaining smooth and natural co-speech gesture generation and preserving the distinct characteristics of the speaker, thereby outperforming state-of-the-art baselines.

手势生成扩散模型风格保留语义控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。