一拍即合:1秒生成高保真表情视频,速度超10倍
Real-time One-Step Diffusion-based Expressive Portrait Videos Generation
- 仅用一步采样实现高保真人脸视频生成
- 相比现有方法提升10倍以上速度,保持同等画质
- 适合实时互动、直播及低延迟应用场景
潜在扩散模型在单参考图像和音频输入下已能生成精准口型同步与自然动作的表达性头像视频。然而,这些模型远未达到实时要求,通常需数十步采样,生成1秒视频耗时数分钟,严重限制实际应用。我们提出OSA-LCM(One-Step Avatar Latent Consistency Model),首次实现基于扩散模型的实时头像视频生成。该方法仅需一步采样,即可达到与现有方法相当的视频质量,速度提升超过10倍。为此,我们设计新型头像判别器,引导口型-音频一致性与动作表现力,在有限采样步骤下提升视频质量。此外,采用两阶段训练架构,结合编辑微调方法(EFT),将生成过程转化为训练阶段的编辑任务,有效缓解单步生成中的时间间隙问题。实验表明,OSA-LCM在开放源代码头像视频生成模型中表现更优,且以单步采样高效运行。
原文摘要 · Abstract (English)
Latent diffusion models have made great strides in generating expressive portrait videos with accurate lip-sync and natural motion from a single reference image and audio input. However, these models are far from real-time, often requiring many sampling steps that take minutes to generate even one second of video-significantly limiting practical use. We introduce OSA-LCM (One-Step Avatar Latent Consistency Model), paving the way for real-time diffusion-based avatars. Our method achieves comparable video quality to existing methods but requires only one sampling step, making it more than 10x faster. To accomplish this, we propose a novel avatar discriminator design that guides lip-audio consistency and motion expressiveness to enhance video quality in limited sampling steps. Additionally, we employ a second-stage training architecture using an editing fine-tuned method (EFT), transforming video generation into an editing task during training to effectively address the temporal gap challenge in single-step generation. Experiments demonstrate that OSA-LCM outperforms existing open-source portrait video generation models while operating more efficiently with a single sampling step.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。