arXiv:2508.17404cs.CV2025-08中稿 · ICLR被引 3

通过解耦结构与外观,生成更自然的人体动作视频。

MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling

  • 分两阶段生成:先建模人体运动结构,再合成外观细节。
  • 在多个指标上显著优于现有方法,尤其在复杂动作和交互上表现突出。
  • 适合需要精准人体动作控制的视频生成、动画设计场景。

现有视频生成模型多关注外观保真度,但在生成复杂人体动作(如全身运动、长时动态、精细人-环境交互)方面能力有限,常导致动作不真实或物理上不合理。为此,我们提出MoSA,将人体视频生成过程解耦为结构生成与外观生成两部分:首先通过3D结构变换器从文本提示生成人体运动序列,随后在该结构引导下合成视频外观。通过引入人体感知动态控制模块并施加密集追踪约束,实现对稀疏人体结构的细粒度控制;通过接触约束提升人-环境交互建模能力。两项机制协同保障生成视频在结构与外观上的保真性。本文还构建了一个大规模人体视频数据集,其动作复杂度和多样性超过现有数据集。大量对比实验表明,MoSA在多数评估指标上显著优于各类通用视频生成、人体视频生成及人体动画模型。

原文摘要 · Abstract (English)

Existing video generation models predominantly emphasize appearance fidelity while exhibiting limited ability to synthesize complex human motions, such as whole-body movements, long-range dynamics, and fine-grained human-environment interactions. This often leads to unrealistic or physically implausible movements with inadequate structural coherence. To conquer these challenges, we propose MoSA, which decouples the process of human video generation into two components, i.e., structure generation and appearance generation. MoSA first employs a 3D structure transformer to generate a human motion sequence from the text prompt. The remaining video appearance is then synthesized under the guidance of this structural sequence. We achieve fine-grained control over the sparse human structures by introducing Human-Aware Dynamic Control modules with a dense tracking constraint during training. The modeling of human-environment interactions is improved through the proposed contact constraint. Those two components work comprehensively to ensure the structural and appearance fidelity across the generated videos. This paper also contributes a large-scale human video dataset, which features more complex and diverse motions than existing human video datasets. We conduct comprehensive comparisons between MoSA and a variety of approaches, including general video generation models, human video generation models, and human animation models. Experiments demonstrate that MoSA substantially outperforms existing approaches across the majority of evaluation metrics.

视频生成人体动作结构解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。