不用逐步生成,一次搞定真人头像视频,还更快更准。
DAWN: Dynamic Frame Avatar with Non-autoregressive Diffusion Framework for Talking Head Video Generation
- 用非自回归扩散模型一次性生成整段视频
- 口型同步精准,动作自然,长视频稳定输出
- 适合需要高速生成高质量人脸视频的场景
说话头视频生成旨在从单张人像和语音音频生成生动逼真的说话头视频。尽管基于扩散模型的方法已取得显著进展,但几乎所有方法都依赖自回归策略,存在上下文利用不足、误差累积和生成速度慢的问题。为此,我们提出DAWN(Dynamic frame Avatar With Non-autoregressive diffusion),一种支持全序列一次性生成动态长度视频的框架。该框架包含两个核心组件:(1) 在潜运动空间中生成音频驱动的整体面部动态;(2) 音频驱动的头部姿态与眨眼生成。大量实验表明,该方法能生成真实感强、口型精准、姿态与眨眼自然的视频。同时,具备高生成速度和强外推能力,可稳定生成高质量长视频。这些结果展示了DAWN在说话头视频生成领域的巨大潜力。我们希望该工作能推动扩散模型中非自回归方法的进一步探索。代码将公开于https://github.com/Hanbo-Cheng/DAWN-pytorch。
原文摘要 · Abstract (English)
Talking head generation intends to produce vivid and realistic talking head videos from a single portrait and speech audio clip. Although significant progress has been made in diffusion-based talking head generation, almost all methods rely on autoregressive strategies, which suffer from limited context utilization beyond the current generation step, error accumulation, and slower generation speed. To address these challenges, we present DAWN (Dynamic frame Avatar With Non-autoregressive diffusion), a framework that enables all-at-once generation of dynamic-length video sequences. Specifically, it consists of two main components: (1) audio-driven holistic facial dynamics generation in the latent motion space, and (2) audio-driven head pose and blink generation. Extensive experiments demonstrate that our method generates authentic and vivid videos with precise lip motions, and natural pose/blink movements. Additionally, with a high generation speed, DAWN possesses strong extrapolation capabilities, ensuring the stable production of high-quality long videos. These results highlight the considerable promise and potential impact of DAWN in the field of talking head video generation. Furthermore, we hope that DAWN sparks further exploration of non-autoregressive approaches in diffusion models. Our code will be publicly available at https://github.com/Hanbo-Cheng/DAWN-pytorch.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。