让人物动画保持身份一致,生成高清视频无需后处理
StableAnimator: High-Quality Identity-Preserving Human Image Animation

- 用参考图和姿态序列端到端生成动画,保持身份不变
- 通过新设计的适配器和优化方法,显著提升人脸质量与一致性
- 适合需要高保真人物动画的影视、虚拟主播等场景
当前的人像动画扩散模型难以保证身份一致性。本文提出 StableAnimator,首个端到端的身份保持视频扩散框架,仅需参考图像和姿态序列即可生成高质量视频,无需任何后处理。基于视频扩散模型,StableAnimator 在训练和推理阶段均设计了专用模块以增强身份一致性:首先使用现成提取器获取图像和人脸嵌入,再通过全局内容感知人脸编码器使二者交互优化;随后引入新型分布感知身份适配器,在避免时序层干扰的同时实现身份对齐。推理阶段提出基于汉密尔顿-雅可比-贝尔曼(HJB)方程的优化方法,将其融入扩散去噪过程,约束去噪路径从而提升身份保留效果。在多个基准测试中,StableAnimator 在定性和定量上均表现出色。
原文摘要 · Abstract (English)
Current diffusion models for human image animation struggle to ensure identity (ID) consistency. This paper presents StableAnimator, the first end-to-end ID-preserving video diffusion framework, which synthesizes high-quality videos without any post-processing, conditioned on a reference image and a sequence of poses. Building upon a video diffusion model, StableAnimator contains carefully designed modules for both training and inference striving for identity consistency. In particular, StableAnimator begins by computing image and face embeddings with off-the-shelf extractors, respectively and face embeddings are further refined by interacting with image embeddings using a global content-aware Face Encoder. Then, StableAnimator introduces a novel distribution-aware ID Adapter that prevents interference caused by temporal layers while preserving ID via alignment. During inference, we propose a novel Hamilton-Jacobi-Bellman (HJB) equation-based optimization to further enhance the face quality. We demonstrate that solving the HJB equation can be integrated into the diffusion denoising process, and the resulting solution constrains the denoising path and thus benefits ID preservation. Experiments on multiple benchmarks show the effectiveness of StableAnimator both qualitatively and quantitatively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。