arXiv:2409.07649cs.CV2024-09CVPR被引 27

用扩散模型生成自然多样的讲话手势,实现单图驱动的TED视频自动生成

DiffTED: One-shot Audio-driven TED Talk Video Generation with Diffusion-based Co-speech Gestures

论文配图:DiffTED: One-shot Audio-driven TED Talk Video Generation with Diffusion-based Co-speech Gestures
图 1 · 摘自论文原文
  • 用扩散模型生成关键点序列,驱动面部与手势同步动画
  • 无需预训练分类器,通过无分类器引导实现手势与音频自然同步
  • 支持单张图像输入,生成时序连贯且多样化的演讲视频

音频驱动的说话人视频生成已取得显著进展,但现有方法通常依赖视频到视频的转换技术及传统的生成网络(如GAN),且常将人脸动作与伴随手势分开生成,导致输出不够连贯。此外,这些方法生成的手势往往过于平滑或单调,缺乏多样性,许多以手势为中心的方法也未整合说话头生成。为解决这些问题,我们提出DiffTED,一种从单张图像出发的一次性音频驱动TED风格说话视频生成新方法。具体而言,我们利用扩散模型生成薄板样条运动模型的关键点序列,精确控制虚拟角色动画,同时确保时序上连贯且多样化的手势。该创新方法采用无分类器引导,使手势能自然随音频输入流动,无需依赖预训练分类器。实验表明,DiffTED可生成时序连贯、手势多样的说话视频。

原文摘要 · Abstract (English)

Audio-driven talking video generation has advanced significantly, but existing methods often depend on video-to-video translation techniques and traditional generative networks like GANs and they typically generate taking heads and co-speech gestures separately, leading to less coherent outputs. Furthermore, the gestures produced by these methods often appear overly smooth or subdued, lacking in diversity, and many gesture-centric approaches do not integrate talking head generation. To address these limitations, we introduce DiffTED, a new approach for one-shot audio-driven TED-style talking video generation from a single image. Specifically, we leverage a diffusion model to generate sequences of keypoints for a Thin-Plate Spline motion model, precisely controlling the avatar's animation while ensuring temporally coherent and diverse gestures. This innovative approach utilizes classifier-free guidance, empowering the gestures to flow naturally with the audio input without relying on pre-trained classifiers. Experiments demonstrate that DiffTED generates temporally coherent talking videos with diverse co-speech gestures.

视频生成扩散模型语音驱动手势生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。