arXiv:2505.10238cs.CV2025-05被引 2

直接用4D动作序列生成角色动画,支持任意角色零样本泛化。

MTVCraft: Tokenizing 4D Motion for Arbitrary Character Animation

  • 将3D动作序列转为4D动作令牌,保留时空信息
  • 在TikTok和Fashion数据集上达到顶尖性能
  • 可零样本控制全身/半身角色及非人类物体

角色图像动画随着数字人兴起迅速发展。现有方法多依赖2D姿态图进行运动引导,限制了泛化能力并丢失了开放世界动画所需的4D信息。为此,我们提出首个直接建模原始3D运动序列(即4D运动)的框架——MTVCraft。核心是4DMoT(4D动作分词器),将3D运动序列量化为4D动作令牌,相比2D姿态图提供更鲁棒的时空线索,且无需严格对齐像素,实现更灵活的解耦控制。接着引入MV-DiT(运动感知视频DiT),通过设计含4D位置编码的独特运动注意力机制,有效利用动作令牌作为紧凑而丰富的4D上下文,指导角色图像在复杂4D世界中的动画生成。我们在CogVideoX-5B(小规模)与Wan-2.1-14B(大规模)上实现该框架,证明其可扩展性。在TikTok与Fashion基准测试中表现达当前最优。凭借鲁棒的动作令牌,MTVCraft展现出前所未有的零样本泛化能力,能动画化任意角色(全身体/半身体)及跨风格、场景的非人类物体。这标志着该领域的重要进展,并开辟了基于姿态引导视频生成的新方向。项目主页:https://github.com/DINGYANB/MTVCrafter;简化版已商用,可通过 https://telestudio.teleagi.cn/generatevideo/creativeWorkshop 使用。

原文摘要 · Abstract (English)

Character image animation has rapidly advanced with the rise of digital humans. However, existing methods rely largely on 2D-rendered pose images for motion guidance, which limits generalization and discards essential 4D information for open-world animation. To address this, we propose MTVCraft (Motion Tokenization Video Crafter), the first framework that directly models raw 3D motion sequences (i.e., 4D motion) for character image animation. Specifically, we introduce 4DMoT (4D motion tokenizer) to quantize 3D motion sequences into 4D motion tokens. Compared to 2D-rendered pose images, 4D motion tokens offer more robust spatial-temporal cues and avoid strict pixel-level alignment between pose images and the character, enabling more flexible and disentangled control. Next, we introduce MV-DiT (Motion-aware Video DiT). By designing unique motion attention with 4D positional encodings, MV-DiT can effectively leverage motion tokens as 4D compact yet expressive context for character image animation in the complex 4D world. We implement MTVCraft on both CogVideoX-5B (small scale) and Wan-2.1-14B (large scale), demonstrating that our framework is easily scalable and can be applied to models of varying sizes. Experiments on the TikTok and Fashion benchmarks demonstrate our state-of-the-art performance. Moreover, powered by robust motion tokens, MTVCraft showcases unparalleled zero-shot generalization. It can animate arbitrary characters in full-body and half-body forms, and even non-human objects across diverse styles and scenarios. Hence, it marks a significant step forward in this field and opens a new direction for pose-guided video generation. Our project page is available at https://github.com/DINGYANB/MTVCrafter. A scaled version has been commercially deployed and is available at https://telestudio.teleagi.cn/generatevideo/creativeWorkshop.

动作生成4D建模零样本视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。