让语音驱动的虚拟人生成更快更稳,速度提升近13倍。
FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation
- 用混合监督损失融合不同质量数据,提升模型鲁棒性。
- 多条件蒸馏加可学习标记,推理速度提升4.17至12.5倍。
- 适合需要高速高质量虚拟人生成的应用场景。
基于扩散模型的语音驱动说话头像方法因高保真、生动且富有表现力而受到关注,但其推理速度慢限制了实际应用。尽管已有多种扩散模型蒸馏技术,我们发现直接蒸馏效果不佳:蒸馏模型对开放集输入图像鲁棒性下降,且音视频相关性低于教师模型,削弱了扩散模型的优势。为此,我们提出FADA(Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation)。首先设计混合监督损失,利用不同质量数据提升整体能力与鲁棒性;其次提出带可学习标记的多条件蒸馏,利用音频与参考图像条件间的关联,将多条件蒸馏导致的三重推理降至单次,仅带来可接受的质量下降。在多个数据集上的实验表明,FADA生成的视频生动逼真,媲美最新扩散模型方法,同时实现NFE速度提升4.17至12.5倍。演示视频见http://fadavatar.github.io。
原文摘要 · Abstract (English)
Diffusion-based audio-driven talking avatar methods have recently gained attention for their high-fidelity, vivid, and expressive results. However, their slow inference speed limits practical applications. Despite the development of various distillation techniques for diffusion models, we found that naive diffusion distillation methods do not yield satisfactory results. Distilled models exhibit reduced robustness with open-set input images and a decreased correlation between audio and video compared to teacher models, undermining the advantages of diffusion models. To address this, we propose FADA (Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation). We first designed a mixed-supervised loss to leverage data of varying quality and enhance the overall model capability as well as robustness. Additionally, we propose a multi-CFG distillation with learnable tokens to utilize the correlation between audio and reference image conditions, reducing the threefold inference runs caused by multi-CFG with acceptable quality degradation. Extensive experiments across multiple datasets show that FADA generates vivid videos comparable to recent diffusion model-based methods while achieving an NFE speedup of 4.17-12.5 times. Demos are available at our webpage http://fadavatar.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。