让人脸动画实时生成,直播可用
PersonaLive! Expressive Portrait Image Animation for Live Streaming
- 用隐式表情和3D关键点控制动作,表达更自然
- 减少推理步数,速度比之前快7到22倍
- 支持低延迟长视频流,适合直播场景
当前基于扩散模型的人脸动画方法主要关注视觉质量和表情真实感,却忽视了生成延迟与实时性能,限制了其在直播场景的应用。我们提出PersonaLive,一种面向实时直播的人脸动画扩散框架,采用多阶段训练策略。首先,结合隐式面部表示与3D隐式关键点,实现精细化图像级运动控制;其次,提出少步数外观蒸馏策略,消除去噪过程中的外观冗余,显著提升推理效率;最后,引入自回归微块流式生成范式,配合滑动训练与历史关键帧机制,实现低延迟、稳定的长期视频生成。大量实验表明,PersonaLive在保持领先性能的同时,相比先前扩散模型实现最高达7-22倍的速度提升。
原文摘要 · Abstract (English)
Current diffusion-based portrait animation models predominantly focus on enhancing visual quality and expression realism, while overlooking generation latency and real-time performance, which restricts their application range in the live streaming scenario. We propose PersonaLive, a novel diffusion-based framework towards streaming real-time portrait animation with multi-stage training recipes. Specifically, we first adopt hybrid implicit signals, namely implicit facial representations and 3D implicit keypoints, to achieve expressive image-level motion control. Then, a fewer-step appearance distillation strategy is proposed to eliminate appearance redundancy in the denoising process, greatly improving inference efficiency. Finally, we introduce an autoregressive micro-chunk streaming generation paradigm equipped with a sliding training strategy and a historical keyframe mechanism to enable low-latency and stable long-term video generation. Extensive experiments demonstrate that PersonaLive achieves state-of-the-art performance with up to 7-22x speedup over prior diffusion-based portrait animation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。