用少步数实现高质量人体图像动画,解决加速导致的模糊失真问题。
Taming Consistency Distillation for Accelerated Human Image Animation
- 分段一致性蒸馏+轻量头,减少累积误差
- 2-4步即可达顶尖模型质量,运动更连贯
- 适合需要快速生成高清人物动画的场景
近期的人体图像动画进展主要依赖视频扩散模型,但其需大量迭代去噪步骤,导致推理成本高、速度慢。采用一致性模型通过一致性蒸馏可有效加速,但在人体动画中常引发视觉模糊、运动退化和面部失真,尤其在动态区域。本文提出DanceLCM方法,结合多项改进:(1) 分段一致性蒸馏配合辅助轻量头,利用真实视频潜在表示提供监督,缓解单轨迹生成带来的累积误差;(2) 引入运动聚焦损失,强化运动区域建模,并显式注入面部保真特征以提升面部真实性。大量定性与定量实验表明,DanceLCM仅需2-4次推理步骤即可达到当前最优视频扩散模型的效果,显著降低推理开销且不牺牲视频质量。代码与模型将公开发布。
原文摘要 · Abstract (English)
Recent advancements in human image animation have been propelled by video diffusion models, yet their reliance on numerous iterative denoising steps results in high inference costs and slow speeds. An intuitive solution involves adopting consistency models, which serve as an effective acceleration paradigm through consistency distillation. However, simply employing this strategy in human image animation often leads to quality decline, including visual blurring, motion degradation, and facial distortion, particularly in dynamic regions. In this paper, we propose the DanceLCM approach complemented by several enhancements to improve visual quality and motion continuity at low-step regime: (1) segmented consistency distillation with an auxiliary light-weight head to incorporate supervision from real video latents, mitigating cumulative errors resulting from single full-trajectory generation; (2) a motion-focused loss to centre on motion regions, and explicit injection of facial fidelity features to improve face authenticity. Extensive qualitative and quantitative experiments demonstrate that DanceLCM achieves results comparable to state-of-the-art video diffusion models with a mere 2-4 inference steps, significantly reducing the inference burden without compromising video quality. The code and models will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。