无需对齐参考图即可实现高质量角色动画与姿态迁移
One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer
- 将训练转为自监督补全任务,处理任意布局的参考图像
- 支持部分可见参考,能提取完整身份特征并融合多分辨率输入
- 通过解耦姿态与外观提升生成质量,适合长视频生成
扩散模型的进展显著提升了姿态驱动的角色动画效果。然而,现有方法受限于空间对齐的参考姿态对和匹配的骨骼结构,无法处理参考姿态错位问题。为此,我们提出 One-to-All Animation,一个统一框架,可在任意布局的参考图像上实现高保真角色动画与图像姿态迁移。首先,为处理空间错位的参考,我们将训练重构为自监督补全任务,将多样布局的参考转换为统一的遮挡输入格式。其次,为处理部分可见的参考,设计参考特征提取器以全面提取身份特征。进一步引入混合参考融合注意力机制,支持不同分辨率和动态序列长度。最后,从生成质量角度,提出身份鲁棒的姿态控制,解耦外观与骨骼结构以缓解姿态过拟合,并采用令牌替换策略保证长视频生成连贯性。大量实验表明,该方法优于现有方法。代码与模型已开源。
原文摘要 · Abstract (English)
Recent advances in diffusion models have greatly improved pose-driven character animation. However, existing methods are limited to spatially aligned reference-pose pairs with matched skeletal structures. Handling reference-pose misalignment remains unsolved. To address this, we present One-to-All Animation, a unified framework for high-fidelity character animation and image pose transfer for references with arbitrary layouts. First, to handle spatially misaligned reference, we reformulate training as a self-supervised outpainting task that transforms diverse-layout reference into a unified occluded-input format. Second, to process partially visible reference, we design a reference extractor for comprehensive identity feature extraction. Further, we integrate hybrid reference fusion attention to handle varying resolutions and dynamic sequence lengths. Finally, from the perspective of generation quality, we introduce identity-robust pose control that decouples appearance from skeletal structure to mitigate pose overfitting, and a token replace strategy for coherent long-video generation. Extensive experiments show that our method outperforms existing approaches. The code and model are available at https://github.com/ssj9596/One-to-All-Animation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。