arXiv:2411.12831cs.CV2024-11ECCV

用文本生成视频模型驱动人体动作,实现真实运动合成。

Towards motion from video diffusion models

  • 通过SMPL-X人体模型与视频扩散模型结合,用分数蒸馏采样生成动作。
  • 在公开文本到视频模型上验证,可生成多样且合理的肢体运动。
  • 为角色动画和动作生成提供新思路,适合动画与虚拟人研究者。

文本条件的视频扩散模型已成为视频生成与编辑的强大工具,但其对人类动作细节的捕捉能力仍待探索。这类模型对多种文本提示的忠实建模,为人物与角色动画带来广泛应用前景。本文初步探讨这些模型能否有效引导真实人体动作的合成。我们提出一种方法:基于视频扩散模型计算得分蒸馏采样(SDS),对SMPL-X人体表示进行形变,以生成动作。通过分析生成动画的保真度,我们评估了利用公开文本到视频扩散模型结合SDS生成运动的能力。研究结果揭示了此类模型在生成多样化、合理人体动作方面的潜力与局限,为该领域的进一步研究奠定了基础。

原文摘要 · Abstract (English)

Text-conditioned video diffusion models have emerged as a powerful tool in the realm of video generation and editing. But their ability to capture the nuances of human movement remains under-explored. Indeed the ability of these models to faithfully model an array of text prompts can lead to a wide host of applications in human and character animation. In this work, we take initial steps to investigate whether these models can effectively guide the synthesis of realistic human body animations. Specifically we propose to synthesize human motion by deforming an SMPL-X body representation guided by Score distillation sampling (SDS) calculated using a video diffusion model. By analyzing the fidelity of the resulting animations, we gain insights into the extent to which we can obtain motion using publicly available text-to-video diffusion models using SDS. Our findings shed light on the potential and limitations of these models for generating diverse and plausible human motions, paving the way for further research in this exciting area.

动作生成扩散模型人体建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。