解决多人动画中身份混淆问题,实现任意数量角色的精准动作绑定。
AnyCrowd: Instance-Isolated Identity-Pose Binding for Arbitrary Multi-Character Animation
- 通过独立编码角色潜变量,避免身份信息纠缠。
- 提出三阶段解耦注意力机制,精确绑定身份与动作序列。
- 适合需要高精度多人角色动画生成的研究与应用。
可控角色动画近年来发展迅速,但多角色动画仍缺乏深入探索。随着角色数量增加,多角色参考编码易出现潜在身份纠缠,导致身份泄露和控制力下降。同时,学习参考身份与驱动动作序列间精确且时空一致的对应关系日益困难,常引发身份-动作错配及生成视频不一致。为此,我们提出 AnyCrowd,一种基于扩散变换器(DiT)的视频生成框架,可扩展至任意数量角色。首先引入实例隔离潜变量表示(IILR),在 DiT 处理前独立编码角色实例,防止潜空间身份纠缠。在此解耦表示基础上,进一步提出三阶段解耦注意力(TSDA),将自注意力分解为:(i) 实例感知前景注意力,(ii) 背景中心交互,(iii) 全局前景-背景协调。此外,为缓解重叠区域的标记歧义,TSDA 内集成自适应门控融合(AGF)模块,预测身份感知权重,有效融合竞争性标记组,形成身份一致的表示。
原文摘要 · Abstract (English)
Controllable character animation has advanced rapidly in recent years, yet multi-character animation remains underexplored. As the number of characters grows, multi-character reference encoding becomes more susceptible to latent identity entanglement, resulting in identity bleeding and reduced controllability. Moreover, learning precise and spatio-temporally consistent correspondences between reference identities and driving pose sequences becomes increasingly challenging, often leading to identity-pose mis-binding and inconsistency in generated videos. To address these challenges, we propose AnyCrowd, a Diffusion Transformer (DiT)-based video generation framework capable of scaling to an arbitrary number of characters. Specifically, we first introduce an Instance-Isolated Latent Representation (IILR), which encodes character instances independently prior to DiT processing to prevent latent identity entanglement. Building on this disentangled representation, we further propose Tri-Stage Decoupled Attention (TSDA) to bind identities to driving poses by decomposing self-attention into: (i) instance-aware foreground attention, (ii) background-centric interaction, and (iii) global foreground-background coordination. Furthermore, to mitigate token ambiguity in overlapping regions, an Adaptive Gated Fusion (AGF) module is integrated within TSDA to predict identity-aware weights, effectively fusing competing token groups into identity-consistent representations...
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。