arXiv:2602.21581cs.CV2026-02中稿 · ed被引 4

让多人图像动画更真实,解决身份混淆和遮挡问题。

MultiAnimate: Pose-Guided Image Animation Made Extensible

  • 用标识分配器与适配器捕捉个体位置和空间关系
  • 仅在两人数据上训练,可泛化到更多角色场景
  • 支持单人与多人动画统一处理,灵活扩展

姿态引导的人体图像动画旨在合成由姿态序列驱动的参考角色真实视频。尽管基于扩散的方法已取得显著进展,但现有方法大多局限于单角色动画。我们观察到,将这些方法直接扩展至多角色场景常导致身份混淆和角色间不合理的遮挡。为此,本文提出一种基于现代扩散变换器(DiTs)的可扩展多角色图像动画框架。核心引入两个新组件——标识分配器与标识适配器,协同捕捉个体位置信息及角色间的空间关系。该掩码驱动机制结合可扩展训练策略,不仅提升灵活性,还使模型能泛化至训练时未见的角色数量。值得注意的是,仅在两人数据集上训练的模型即可实现多角色动画,同时保持与单角色场景的兼容性。大量实验表明,该方法在多角色图像动画任务中达到领先性能,优于现有扩散基基线。

原文摘要 · Abstract (English)

Pose-guided human image animation aims to synthesize realistic videos of a reference character driven by a sequence of poses. While diffusion-based methods have achieved remarkable success, most existing approaches are limited to single-character animation. We observe that naively extending these methods to multi-character scenarios often leads to identity confusion and implausible occlusions between characters. To address these challenges, in this paper, we propose an extensible multi-character image animation framework built upon modern Diffusion Transformers (DiTs) for video generation. At its core, our framework introduces two novel components-Identifier Assigner and Identifier Adapter - which collaboratively capture per-person positional cues and inter-person spatial relationships. This mask-driven scheme, along with a scalable training strategy, not only enhances flexibility but also enables generalization to scenarios with more characters than those seen during training. Remarkably, trained on only a two-character dataset, our model generalizes to multi-character animation while maintaining compatibility with single-character cases. Extensive experiments demonstrate that our approach achieves state-of-the-art performance in multi-character image animation, surpassing existing diffusion-based baselines.

图像动画扩散模型多角色姿态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。