一键替换视频中的人,保持动作和外观自然
Replace Anyone in Videos
- 用姿势引导的图像条件视频修复框架实现局部人物替换
- 支持复杂背景下的无缝替换,保留姿态与外观一致性
- 适用于3D-UNet和DiT模型,适合影视后期与虚拟制作
可控人体中心视频生成领域已取得显著进展,尤其得益于扩散模型的兴起。然而,在视频中精确、局部地控制人体运动,如替换或插入人物并保持期望的动作模式,仍是重大挑战。本文提出ReplaceAnyone框架,专注于具有复杂背景的局部人体替换与插入。我们将该任务建模为带有姿态引导的图像条件视频修复问题,采用统一的端到端视频扩散架构,在掩码区域内实现图像条件视频修复。为防止形体泄露并实现细粒度局部控制,引入多种掩码形式,包括规则与不规则形状。此外,设计增强视觉引导机制以提升外观对齐效果,采用混合修复编码器以更好地保留掩码区域的背景细节,并提出两阶段优化方法以降低训练难度。ReplaceAnyone可在单一框架内实现角色的无缝替换或插入,同时保持目标姿态动作与参考外观。大量实验表明,该方法能生成真实且连贯的视频内容。所提方法可无缝应用于传统3D-UNet基模型及基于DiT的视频模型(如Wan2.1)。代码将公开于https://github.com/ali-vilab/UniAnimate-DiT。
原文摘要 · Abstract (English)
The field of controllable human-centric video generation has witnessed remarkable progress, particularly with the advent of diffusion models. However, achieving precise and localized control over human motion in videos, such as replacing or inserting individuals while preserving desired motion patterns, still remains a formidable challenge. In this work, we present the ReplaceAnyone framework, which focuses on localized human replacement and insertion featuring intricate backgrounds. Specifically, we formulate this task as an image-conditioned video inpainting paradigm with pose guidance, utilizing a unified end-to-end video diffusion architecture that facilitates image-conditioned video inpainting within masked regions. To prevent shape leakage and enable granular local control, we introduce diverse mask forms involving both regular and irregular shapes. Furthermore, we implement an enriched visual guidance mechanism to enhance appearance alignment, a hybrid inpainting encoder to further preserve the detailed background information in the masked video, and a two-phase optimization methodology to simplify the training difficulty. ReplaceAnyone enables seamless replacement or insertion of characters while maintaining the desired pose motion and reference appearance within a single framework. Extensive experimental results demonstrate the effectiveness of our method in generating realistic and coherent video content. The proposed ReplaceAnyone can be seamlessly applied not only to traditional 3D-UNet base models but also to DiT-based video models such as Wan2.1. The code will be available at https://github.com/ali-vilab/UniAnimate-DiT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。