任意角色+场景生成视频,用姿态引导实现灵活可控
AnyCharV: Bootstrap Controllable Character Video Generation with Fine-to-Coarse Guidance
- 两阶段训练:先融合角色与场景,再通过自提升优化细节
- 相比现有方法,角色特征保留更完整,生成视频质量更高
- 适合需要自由定制角色动作和场景的创作者使用
角色视频生成是重要现实应用,旨在生成特定角色的高质量视频。近期方法引入多种控制信号以动画化静态角色,提升了生成过程的可控性。然而,这些方法灵活性不足,难以将源角色合成至目标场景。为此,我们提出AnyCharV框架,可灵活使用任意源角色和目标场景,并在姿态信息引导下生成角色视频。方法采用两阶段训练:第一阶段构建基础模型,实现角色与场景的融合;第二阶段通过自提升机制,用第一阶段生成视频替换细掩码为粗掩码,训练出更注重角色细节保留的结果。大量实验表明,本方法优于现有最先进方法。
原文摘要 · Abstract (English)
Character video generation is a significant real-world application focused on producing high-quality videos featuring specific characters. Recent advancements have introduced various control signals to animate static characters, successfully enhancing control over the generation process. However, these methods often lack flexibility, limiting their applicability and making it challenging for users to synthesize a source character into a desired target scene. To address this issue, we propose a novel framework, AnyCharV, that flexibly generates character videos using arbitrary source characters and target scenes, guided by pose information. Our approach involves a two-stage training process. In the first stage, we develop a base model capable of integrating the source character with the target scene using pose guidance. The second stage further bootstraps controllable generation through a self-boosting mechanism, where we use the generated video in the first stage and replace the fine mask with the coarse one, enabling training outcomes with better preservation of character details. Extensive experimental results demonstrate the superiority of our method compared with previous state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。