端到端动画框架,实现实时角色驱动,支持视角可控
Wan-Animate-2: Pushing the Application Boundaries of Character Animation

- 用扩散变换器直接处理驱动视频,无中间运动提取
- 支持文本控制视角,生成质量高且身份保持稳定
- 轻量版可实时推理,适合直播与数字人等交互场景
角色图像动画是计算机视觉中的基础挑战。现有方法分为三类:基于显式运动表示的方法易出现提取误差和身份漂移;基于隐式运动特征的方法会因压缩丢失细节;基于上下文学习的方法虽避免中间表示但计算成本过高。此外,当前系统均面向离线合成,无法满足数字人、直播主播等交互应用的实时需求。为此,我们提出 Wan-Animate-2,一个端到端的角色动画框架,通过重构的扩散变换器直接输入驱动视频,完全消除中间运动提取器,实现更高运动保真度与身份一致性。我们进一步引入文本驱动的视角控制,解耦输出相机视角与驱动视频,该能力在依赖显式运动表示的方法中极为罕见。在生成质量之外,我们提出 Wan-Animate-2-Lite,通过三阶段训练范式(教师强制预训练+误差缓存机制,自强化蒸馏+分块反向传播)显著降低推理延迟至实时水平,使流式角色动画成为可能,拓展了此前不可行的应用场景。定性评估与用户研究显示,Wan-Animate-2 在多样角色与动作模式下均能生成高保真结果。为推动研究与社区发展,我们将公开 Wan-Animate-2-Base 模型权重。
原文摘要 · Abstract (English)
Character image animation remains a foundational yet challenging task in computer vision. Existing approaches can be broadly categorized into three paradigms: methods based on explicit motion representations suffer from extraction errors and identity drift; methods based on implicit motion features lose fine-grained dynamics through compression; and in-context learning approaches avoid intermediate representations but incur prohibitive computational costs. Furthermore, all current systems are designed for offline synthesis, unable to meet the real-time requirements of interactive applications such as digital avatars and live-streaming hosts. To address these limitations, we present Wan-Animate-2, an end-to-end character animation framework that directly consumes the driving video within a redesigned Diffusion Transformer. Our architecture achieves superior motion fidelity and identity preservation by eliminating intermediate motion extractors entirely. We further introduce text driven viewpoint control that decouples the output camera perspective from the driving video--a capability rarely supported by prior character animation methods that rely on explicit motion representations. Beyond generation quality, we present Wan-Animate-2-Lite, an efficient variant that reduces inference latency to real-time thresholds through a three-stage training paradigm: teacher forcing pretraining with error buffer mechanism, and Self-Forcing distillation with chunk-wise backpropagation. This enables streaming character animation for interactive applications, opening new deployment scenarios that were previously infeasible. Qualitative evaluations and user studies demonstrate that Wan-Animate-2 achieves high-fidelity animation results across diverse characters and motion patterns. To foster further research and community development, we will release the Wan-Animate-2-Base model weights to the public.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。