用扩散模型一步生成自然连贯的口型同步手势视频,数据需求少。
EasyGenNet: An Efficient Framework for Audio-Driven Gesture Video Generation Based on Diffusion Model

- 基于扩散模型的一阶段训练与时间推理,无需额外时序模块训练。
- 仅需数千帧数据即可完成角色微调,显著降低数据依赖。
- 以2D骨骼为中间表示,生成效果优于现有GAN和扩散模型方法。
语音驱动的口型同步视频生成通常分为两步:语音转手势、手势转视频。尽管语音转手势已取得进展,但手势转视频系统在生成自然表情和动作方面仍具挑战性。以往方法采用复杂输入与训练策略,并需大规模数据集预训练,限制了实际应用。本文提出一种基于扩散模型的简单一阶段训练方法与时间推理机制,无需额外时序模块训练即可生成真实且连续的手势视频。整个模型复用现有预训练权重,每个角色仅需数千帧数据即可完成微调。基于视频生成器,我们构建新的音频到视频管道,使用2D人体骨骼作为中间运动表示。实验表明,该方法在生成质量上优于现有基于GAN和扩散模型的方法。
原文摘要 · Abstract (English)
Audio-driven cospeech video generation typically involves two stages: speech-to-gesture and gesture-to-video. While significant advances have been made in speech-to-gesture generation, synthesizing natural expressions and gestures remains challenging in gesture-to-video systems. In order to improve the generation effect, previous works adopted complex input and training strategies and required a large amount of data sets for pre-training, which brought inconvenience to practical applications. We propose a simple one-stage training method and a temporal inference method based on a diffusion model to synthesize realistic and continuous gesture videos without the need for additional training of temporal modules.The entire model makes use of existing pre-trained weights, and only a few thousand frames of data are needed for each character at a time to complete fine-tuning. Built upon the video generator, we introduce a new audio-to-video pipeline to synthesize co-speech videos, using 2D human skeleton as the intermediate motion representation. Our experiments show that our method outperforms existing GAN-based and diffusion-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。