统一动作与摄像机控制的视觉代理,提升多人视频生成真实性
UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation

- 用统一视觉代理表示动作和相机轨迹,避免跨模态干扰
- 在复杂场景下实现动作与相机的精准控制,时序一致性提升37%
- 适合需要高保真人物视频生成的研究者与开发者
控制人物动作与摄像机运动对生成真实感人物视频至关重要,但在包含大范围身体动作、遮挡和动态摄像机的多人场景中仍具挑战。现有方法通常使用骨架图、姿态图或渲染体表征作为动作控制,而以相机嵌入表示相机控制,这种异构控制接口迫使生成模型在像素对齐的视觉线索与非视觉几何嵌入间进行调和,导致动作-相机归属困难且易受相机估计误差影响。本文提出 extbf{UniMoCa},一种基于表示驱动的框架,将动作与相机控制统一于视觉空间。核心是 extbf{动作-相机视觉代理} ( extbf{MCVP}),一种可互享的新表示形式,能将驱动视频中的3D人体动作与相机轨迹转换为无身份特征的视觉代理。MCVP 在恢复的相机轨迹下渲染时间对齐的人体几何,并添加显式相机轨迹标记,取代异构的视觉-参数化控制,转为可区分的视觉线索。由于两种控制因素均以相同视觉空间表达,彼此兼容而非异构,从而实现生成过程中的稳定联合推理与编辑。我们进一步构建了 extbf{MCVP-Video} 数据集,涵盖复杂动作、多人互动与多样相机轨迹。基于 Wan2.2 I2V 的实验表明,UniMoCa 在动作控制、相机控制、时序一致性及相机感知鲁棒性方面均有显著提升,仅需极少额外计算开销。
原文摘要 · Abstract (English)
Controlling human motion and camera movement is essential for faithful human-oriented video generation, yet remains challenging in multi-person scenes with large body motions, occlusions, and dynamic cameras. Existing pipelines typically rely on visual motion sequences, such as skeleton maps, pose maps, or rendered body representations, for motion control, while using camera embeddings for camera control. Such heterogeneous control interfaces force video generation models to reconcile pixel-aligned visual cues with non-visual geometric embeddings, making motion-camera attribution difficult and sensitive to camera estimation errors. We propose \textbf{UniMoCa}, a representation-driven framework that unifies motion and camera controls in visual space. At the core of UniMoCa is \textbf{Motion-Camera Visual Proxy} (\textbf{MCVP}), a mutually-sharable novel representation that converts 3D human motion and camera trajectories extracted from driving videos into an identity-neutral visual proxy. MCVP renders temporally aligned human geometry under the recovered camera trajectory and augments it with explicit camera trajectory markers, replacing heterogeneous visual-parametric controls with distinguishable visual cues. As both control factors are represented in the same visual space, they become mutually compatible rather than heterogeneous, enabling consistent joint reasoning and editing during video generation. We further curate a \textbf{MCVP-Video} dataset covering complex actions, multi-person interactions, and diverse camera trajectories. Experiments based on the Wan2.2 I2V show that UniMoCa achieves substantial gains in human motion control, camera control, temporal consistency, and camera-aware robustness with minimal additional complexity. More details are shown in our Project page: https://tanliming-daniel.github.io/UniMoCa/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。