6倍加速无限长人脸动画,保持身份一致
FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent Prediction
- 用归一化表情块对齐面部特征与扩散潜变量
- 动态滑窗加权融合,实现长视频无断裂
- 高阶潜变量导数跳过去噪步骤,提速6倍
当前基于扩散模型的长视频人脸动画加速方法难以保证身份一致性。本文提出FlashPortrait,一种端到端视频扩散变压器,可在生成保持身份一致的无限长视频的同时,实现最高6倍的推理速度提升。该方法首先使用现成提取器计算与身份无关的面部表情特征;随后引入归一化面部表情模块,通过均值和方差归一化对齐面部特征与扩散潜变量,增强面部建模的身份稳定性。推理时,FlashPortrait采用动态滑动窗口机制,并在重叠区域进行加权融合,确保长动画中过渡平滑且身份一致。在每个上下文窗口内,基于特定时间步的潜变量变化率及各扩散层间导数大小比,直接利用高阶潜变量导数预测未来时间步的潜变量,跳过多个去噪步骤,实现6倍加速。在多个基准测试上的实验表明,FlashPortrait在定性和定量上均有效。
原文摘要 · Abstract (English)
Current diffusion-based acceleration methods for long-portrait animation struggle to ensure identity (ID) consistency. This paper presents FlashPortrait, an end-to-end video diffusion transformer capable of synthesizing ID-preserving, infinite-length videos while achieving up to 6x acceleration in inference speed. In particular, FlashPortrait begins by computing the identity-agnostic facial expression features with an off-the-shelf extractor. It then introduces a Normalized Facial Expression Block to align facial features with diffusion latents by normalizing them with their respective means and variances, thereby improving identity stability in facial modeling. During inference, FlashPortrait adopts a dynamic sliding-window scheme with weighted blending in overlapping areas, ensuring smooth transitions and ID consistency in long animations. In each context window, based on the latent variation rate at particular timesteps and the derivative magnitude ratio among diffusion layers, FlashPortrait utilizes higher-order latent derivatives at the current timestep to directly predict latents at future timesteps, thereby skipping several denoising steps and achieving 6x speed acceleration. Experiments on benchmarks show the effectiveness of FlashPortrait both qualitatively and quantitatively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。