arXiv:2410.20359cs.SDcs.AI2024-10中稿 · WACV 2025被引 8

用条件GAN提升音频驱动手势生成速度与真实度

Conditional GAN for Enhancing Diffusion Models in Efficient and Authentic Global Gesture Generation from Audios

  • 引入条件GAN捕捉音频信号,实现跨步内多模态去噪匹配
  • 单步去噪可采大噪声值,生成速度提升且动作更自然
  • 适合需要高效真实手势生成的AI游戏与影视制作场景

音频驱动的同步手势生成在人机交互、AI游戏和影视制作中至关重要。现有方法存在局限:基于变分自编码器(VAE)的方法易产生局部抖动与全局不稳定,而基于扩散模型的方法则因生成效率低而受限。这是由于扩散模型中的DDPM假设每一步添加的噪声来自单模分布且数值较小;而DDIM通过借鉴微分方程欧拉法打破马尔可夫链,增大噪声步长以减少去噪步骤,加速生成,但直接增大步长会导致结果偏离原始数据分布,造成动作质量下降与不自然伪影。本文突破DDPM假设,在去噪速度与保真度上取得突破。我们引入条件生成对抗网络(conditional GAN),捕捉音频控制信号,并在同一样本步内隐式匹配扩散与去噪步骤间的多模态去噪分布,实现更大噪声值的采样与更少去噪步数,从而实现高速生成。

原文摘要 · Abstract (English)

Audio-driven simultaneous gesture generation is vital for human-computer communication, AI games, and film production. While previous research has shown promise, there are still limitations. Methods based on VAEs are accompanied by issues of local jitter and global instability, whereas methods based on diffusion models are hampered by low generation efficiency. This is because the denoising process of DDPM in the latter relies on the assumption that the noise added at each step is sampled from a unimodal distribution, and the noise values are small. DDIM borrows the idea from the Euler method for solving differential equations, disrupts the Markov chain process, and increases the noise step size to reduce the number of denoising steps, thereby accelerating generation. However, simply increasing the step size during the step-by-step denoising process causes the results to gradually deviate from the original data distribution, leading to a significant drop in the quality of the generated actions and the emergence of unnatural artifacts. In this paper, we break the assumptions of DDPM and achieves breakthrough progress in denoising speed and fidelity. Specifically, we introduce a conditional GAN to capture audio control signals and implicitly match the multimodal denoising distribution between the diffusion and denoising steps within the same sampling step, aiming to sample larger noise values and apply fewer denoising steps for high-speed generation.

手势生成扩散模型条件GAN音频驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。