arXiv:2508.08588cs.CVeess.IV2025-08被引 9

分离控制人物动作、轨迹、背景和外观,实现任意人物在任意场景中做任意动作的视频生成。

RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space

  • 在3D世界空间中解耦运动与外观、主体与背景、动作与轨迹
  • 通过焦距校准和坐标变换将2D轨迹还原为3D,实现精准路径控制
  • 支持文本生成动作、混合不同元素,适合影视动画与虚拟人应用

生成具有真实且可控制动作的人体视频是一项挑战。现有方法虽能生成视觉吸引人的视频,但难以分别控制前景主体、背景视频、人体运动轨迹和动作模式这四个关键元素。本文提出一种分解式人体运动控制与视频生成框架,显式解耦运动与外观、主体与背景、动作与轨迹,实现这些元素的灵活混搭组合。具体地,我们首先构建基于地面的3D世界坐标系,并在3D空间中直接进行运动编辑;通过焦距校准和坐标变换将编辑后的2D轨迹反投影至3D空间,再进行速度对齐与朝向调整以实现轨迹控制;动作则由动作库提供或通过文本转动作方法生成。随后,基于现代文本到视频扩散变换器模型,我们将主体作为令牌注入全注意力机制,将背景沿通道维度拼接,并通过加法方式注入运动(轨迹与动作)控制信号。该设计使我们能够生成任意人物在任意地点执行任意动作的逼真视频。在基准数据集和真实案例上的大量实验表明,本方法在元素级可控性与整体视频质量上均达到当前最优水平。

原文摘要 · Abstract (English)

Generating human videos with realistic and controllable motions is a challenging task. While existing methods can generate visually compelling videos, they lack separate control over four key video elements: foreground subject, background video, human trajectory and action patterns. In this paper, we propose a decomposed human motion control and video generation framework that explicitly decouples motion from appearance, subject from background, and action from trajectory, enabling flexible mix-and-match composition of these elements. Concretely, we first build a ground-aware 3D world coordinate system and perform motion editing directly in the 3D space. Trajectory control is implemented by unprojecting edited 2D trajectories into 3D with focal-length calibration and coordinate transformation, followed by speed alignment and orientation adjustment; actions are supplied by a motion bank or generated via text-to-motion methods. Then, based on modern text-to-video diffusion transformer models, we inject the subject as tokens for full attention, concatenate the background along the channel dimension, and add motion (trajectory and action) control signals by addition. Such a design opens up the possibility for us to generate realistic videos of anyone doing anything anywhere. Extensive experiments on benchmark datasets and real-world cases demonstrate that our method achieves state-of-the-art performance on both element-wise controllability and overall video quality.

动作控制视频生成3D空间文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。