arXiv:2601.11096cs.CV2026-01被引 5

解决多人动画中角色数量、类型和位置不匹配的问题

CoDance: An Unbind-Rebind Paradigm for Robust Multi-Subject Animation

  • 通过解绑-重绑机制打破姿态与参考图的刚性绑定
  • 支持任意数量角色、类型及空间布局的动画生成
  • 适合需要灵活多角色动画的视频生成场景

角色图像动画在多个领域日益重要,但现有方法在处理多角色、多样化角色类型及参考图与驱动姿态间空间错位时表现不佳。这源于过于刚性的空间绑定和难以准确重绑定运动到目标角色。为此,我们提出CoDance,一种新颖的解绑-重绑框架,可基于单个可能错位的姿态序列,实现任意角色数量、类型和空间配置的动画生成。解绑模块通过新型姿态偏移编码器,在姿态及其潜在特征上引入随机扰动,迫使模型学习无位置依赖的运动表示;重绑模块则利用文本提示的语义引导和角色掩码的空间引导,将学习到的运动精确分配给目标角色。此外,我们还构建了新的多角色评估基准CoDanceBench。在CoDanceBench及现有数据集上的大量实验表明,CoDance达到当前最优性能,展现出对多样化角色和空间布局的强大泛化能力。代码与权重将开源。

原文摘要 · Abstract (English)

Character image animation is gaining significant importance across various domains, driven by the demand for robust and flexible multi-subject rendering. While existing methods excel in single-person animation, they struggle to handle arbitrary subject counts, diverse character types, and spatial misalignment between the reference image and the driving poses. We attribute these limitations to an overly rigid spatial binding that forces strict pixel-wise alignment between the pose and reference, and an inability to consistently rebind motion to intended subjects. To address these challenges, we propose CoDance, a novel Unbind-Rebind framework that enables the animation of arbitrary subject counts, types, and spatial configurations conditioned on a single, potentially misaligned pose sequence. Specifically, the Unbind module employs a novel pose shift encoder to break the rigid spatial binding between the pose and the reference by introducing stochastic perturbations to both poses and their latent features, thereby compelling the model to learn a location-agnostic motion representation. To ensure precise control and subject association, we then devise a Rebind module, leveraging semantic guidance from text prompts and spatial guidance from subject masks to direct the learned motion to intended characters. Furthermore, to facilitate comprehensive evaluation, we introduce a new multi-subject CoDanceBench. Extensive experiments on CoDanceBench and existing datasets show that CoDance achieves SOTA performance, exhibiting remarkable generalization across diverse subjects and spatial layouts. The code and weights will be open-sourced.

角色动画多主体解绑重绑姿态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。