让视频中人物动作与镜头运动分开控制,更灵活精准。
TokenMotion: Decoupled Motion Control via Token Disentanglement for Human-centric Video Generation

- 用时空标记分离表示人物姿态和镜头运动,实现局部精细控制。
- 提出解耦-融合框架,通过人感知动态掩码统一建模两类运动。
- 在文本/图像生成视频任务中表现超越现有方法,适合创意制作场景。
人物中心的视频生成中,同时控制镜头运动与人物姿态仍是一大挑战,尤其在类似格莱美颁奖礼‘闪光机器人’的经典场景中。尽管近期视频扩散模型取得显著进展,但现有方法在运动表征能力有限且难以整合镜头与人物运动控制方面仍存在不足。本文提出 TokenMotion,首个基于 DiT 的视频扩散框架,可对镜头运动、人物运动及其联合交互实现细粒度控制。通过将镜头轨迹与人物姿态表示为时空标记,实现局部控制精度。所提方法采用解耦-融合策略,借助人感知动态掩码,有效处理组合运动信号在空间与时间上的变化特性。大量实验表明,TokenMotion 在文本到视频与图像到视频两种范式下均显著优于当前最先进方法,在人物中心运动控制任务中表现出色。本工作推动了可控视频生成的发展,尤其适用于创意生产应用。
原文摘要 · Abstract (English)
Human-centric motion control in video generation remains a critical challenge, particularly when jointly controlling camera movements and human poses in scenarios like the iconic Grammy Glambot moment. While recent video diffusion models have made significant progress, existing approaches struggle with limited motion representations and inadequate integration of camera and human motion controls. In this work, we present TokenMotion, the first DiT-based video diffusion framework that enables fine-grained control over camera motion, human motion, and their joint interaction. We represent camera trajectories and human poses as spatio-temporal tokens to enable local control granularity. Our approach introduces a unified modeling framework utilizing a decouple-and-fuse strategy, bridged by a human-aware dynamic mask that effectively handles the spatially-and-temporally varying nature of combined motion signals. Through extensive experiments, we demonstrate TokenMotion's effectiveness across both text-to-video and image-to-video paradigms, consistently outperforming current state-of-the-art methods in human-centric motion control tasks. Our work represents a significant advancement in controllable video generation, with particular relevance for creative production applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。