120倍提速生成真人说话视频,且质量不降。
TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation

- 分两阶段渐进式蒸馏,先4步再缩到1步。
- 推理速度提升120倍,画面质量保持高水准。
- 适合需要实时生成的虚拟主播、交互应用。
现有音频驱动视频数字人生成模型依赖多步去噪,计算开销大,严重限制实际部署。尽管单步蒸馏可显著加速推理,但常面临训练不稳定的难题。为此,我们提出TurboTalk,一种两阶段渐进式蒸馏框架,将多步音频驱动视频扩散模型高效压缩为单步生成器。首先采用分布匹配蒸馏获得稳定可靠的4步学生模型,随后通过对抗蒸馏逐步将去噪步骤从4步减少至1步。为确保极端步数缩减下的训练稳定性,引入渐进式时间步采样策略和自对比对抗目标,提供中间对抗参考以稳定蒸馏过程。实验表明,该方法实现单步生成说话头像视频,推理速度提升120倍,同时保持高质量生成效果。
原文摘要 · Abstract (English)
Existing audio-driven video digital human generation models rely on multi-step denoising, resulting in substantial computational overhead that severely limits their deployment in real-world settings. While one-step distillation approaches can significantly accelerate inference, they often suffer from training instability. To address this challenge, we propose TurboTalk, a two-stage progressive distillation framework that effectively compresses a multi-step audio-driven video diffusion model into a single-step generator. We first adopt Distribution Matching Distillation to obtain a strong and stable 4-step student, and then progressively reduce the denoising steps from 4 to 1 through adversarial distillation. To ensure stable training under extreme step reduction, we introduce a progressive timestep sampling strategy and a self-compare adversarial objective that provides an intermediate adversarial reference that stabilizes progressive distillation. Our method achieve single-step generation of video talking avatar, boosting inference speed by 120 times while maintaining high generation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。