首个支持任意角色动作的视频生成框架,可精准控制运动与镜头变化。
PoseAnything: Universal Pose-guided Video Generation with Part-aware Temporal Coherence
- 通用骨骼输入,兼容人体与非人体角色动作控制。
- 分部件时间一致性模块,提升动作连贯性,避免形变失真。
- 独立控制角色与镜头运动,适合动画创作与虚拟拍摄场景。
姿态引导视频生成通过姿态序列控制生成视频中主体的运动,实现对角色动作的精确控制,在动画领域具有重要应用价值。然而,现有方法仅支持人体姿态输入,泛化能力差。为此,我们提出PoseAnything,首个能处理人体与非人体角色的通用姿态引导视频生成框架,支持任意骨骼结构输入。为增强运动过程中的一致性,提出分部件时间一致性模块,将主体划分为多个部分,建立跨帧对应关系,并在对应部分间计算交叉注意力,实现细粒度的部件级一致性。此外,提出主体与相机运动解耦的条件控制生成策略(Subject and Camera Motion Decoupled CFG),首次在姿态引导视频生成中实现独立的相机运动控制,通过将主体与相机运动信息分别注入正负锚点实现。同时,构建了XPose数据集,包含50,000个非人体姿态-视频对,并提供自动化标注与筛选流程。大量实验表明,PoseAnything在效果与泛化能力上均显著优于现有最优方法。
原文摘要 · Abstract (English)
Pose-guided video generation refers to controlling the motion of subjects in generated video through a sequence of poses. It enables precise control over subject motion and has important applications in animation. However, current pose-guided video generation methods are limited to accepting only human poses as input, thus generalizing poorly to pose of other subjects. To address this issue, we propose PoseAnything, the first universal pose-guided video generation framework capable of handling both human and non-human characters, supporting arbitrary skeletal inputs. To enhance consistency preservation during motion, we introduce Part-aware Temporal Coherence Module, which divides the subject into different parts, establishes part correspondences, and computes cross-attention between corresponding parts across frames to achieve fine-grained part-level consistency. Additionally, we propose Subject and Camera Motion Decoupled CFG, a novel guidance strategy that, for the first time, enables independent camera movement control in pose-guided video generation, by separately injecting subject and camera motion control information into the positive and negative anchors of CFG. Furthermore, we present XPose, a high-quality public dataset containing 50,000 non-human pose-video pairs, along with an automated pipeline for annotation and filtering. Extensive experiments demonstrate that Pose-Anything significantly outperforms state-of-the-art methods in both effectiveness and generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。