arXiv:2602.03796cs.CV2026-02被引 3

用隐式3D运动表示实现视角自适应的人体视频生成

3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation

  • 用运动编码器提取视点无关的紧凑运动标记,通过交叉注意力注入生成器
  • 在单视图、多视图和动镜视频上训练,实现跨视角运动一致性
  • 逐步弱化对SMPL的依赖,让模型从数据中学习真实的3D空间运动理解

现有方法在人体视频生成中通常依赖2D姿态或显式的3D参数化模型(如SMPL)作为控制信号。然而,2D姿态将运动严格绑定于驱动视角,无法实现新视角合成;显式3D模型虽具结构信息,但存在深度模糊与动态不准确等固有缺陷,若作为强约束会压制大规模视频生成器的内在3D感知能力。本文从3D感知角度重思运动控制,提出隐式、视点无关的运动表示,自然契合生成器的空间先验。我们引入3DiMo,联合训练运动编码器与预训练视频生成器,将驱动帧提炼为紧凑的视点无关运动标记,并通过交叉注意力语义注入。通过单视图、多视图及动镜视频的丰富视点监督,强制跨视角运动一致性。同时,采用辅助几何监督,仅在早期使用SMPL初始化,随后逐渐衰减至零,促使模型从外部3D引导过渡到从数据与生成器先验中学习真实的3D空间运动理解。实验表明,3DiMo能忠实复现驱动动作,支持灵活的文本驱动相机控制,在运动保真度与视觉质量上显著优于现有方法。

原文摘要 · Abstract (English)

Existing methods for human motion control in video generation typically rely on either 2D poses or explicit 3D parametric models (e.g., SMPL) as control signals. However, 2D poses rigidly bind motion to the driving viewpoint, precluding novel-view synthesis. Explicit 3D models, though structurally informative, suffer from inherent inaccuracies (e.g., depth ambiguity and inaccurate dynamics) which, when used as a strong constraint, override the powerful intrinsic 3D awareness of large-scale video generators. In this work, we revisit motion control from a 3D-aware perspective, advocating for an implicit, view-agnostic motion representation that naturally aligns with the generator's spatial priors rather than depending on externally reconstructed constraints. We introduce 3DiMo, which jointly trains a motion encoder with a pretrained video generator to distill driving frames into compact, view-agnostic motion tokens, injected semantically via cross-attention. To foster 3D awareness, we train with view-rich supervision (i.e., single-view, multi-view, and moving-camera videos), forcing motion consistency across diverse viewpoints. Additionally, we use auxiliary geometric supervision that leverages SMPL only for early initialization and is annealed to zero, enabling the model to transition from external 3D guidance to learning genuine 3D spatial motion understanding from the data and the generator's priors. Experiments confirm that 3DiMo faithfully reproduces driving motions with flexible, text-driven camera control, significantly surpassing existing methods in both motion fidelity and visual quality.

人体生成3D感知运动控制视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。