arXiv:2503.17059cs.GRcs.CV2025-03AAAI被引 4

仅用10步采样实现高质量语音驱动手势实时生成。

DIDiffGes: Decoupled Semi-Implicit Diffusion Models for Real-time Gesture Generation from Speech

  • 解耦身体与手部分布,结合GAN隐式建模与L2显式学习。
  • 10步采样生成效果媲美百步传统方法,质量不降。
  • 适合需要低延迟的虚拟主播、人机交互场景。

扩散模型在生成共言语手势方面展现出卓越的合成质量和多样性,但其计算密集的采样过程限制了实际应用。为此,我们提出DIDiffGes——一种基于解耦半隐式扩散模型的框架,仅需少量采样步骤即可从语音生成高质量、富有表现力的手势。该方法将手势数据解耦为身体与手部分布,并进一步分解为边缘与条件分布;利用生成对抗网络(GAN)隐式建模边缘分布,同时通过L2重建损失显式学习条件分布,从而提升训练稳定性并保障全身手势的表现力。框架还学习基于局部身体表征的根节点去噪,确保生成结果的稳定性和真实感。DIDiffGes仅需10次采样即可生成手势,相比现有方法减少100倍采样步数,且不牺牲质量与表达性。用户研究显示,本方法在逼真度、适当性和风格准确性上均优于当前最优方案。

原文摘要 · Abstract (English)

Diffusion models have demonstrated remarkable synthesis quality and diversity in generating co-speech gestures. However, the computationally intensive sampling steps associated with diffusion models hinder their practicality in real-world applications. Hence, we present DIDiffGes, for a Decoupled Semi-Implicit Diffusion model-based framework, that can synthesize high-quality, expressive gestures from speech using only a few sampling steps. Our approach leverages Generative Adversarial Networks (GANs) to enable large-step sampling for diffusion model. We decouple gesture data into body and hands distributions and further decompose them into marginal and conditional distributions. GANs model the marginal distribution implicitly, while L2 reconstruction loss learns the conditional distributions exciplictly. This strategy enhances GAN training stability and ensures expressiveness of generated full-body gestures. Our framework also learns to denoise root noise conditioned on local body representation, guaranteeing stability and realism. DIDiffGes can generate gestures from speech with just 10 sampling steps, without compromising quality and expressiveness, reducing the number of sampling steps by a factor of 100 compared to existing methods. Our user study reveals that our method outperforms state-of-the-art approaches in human likeness, appropriateness, and style correctness. Project is https://cyk990422.github.io/DIDiffGes.

手势生成扩散模型实时生成语音驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。