arXiv:2410.08931cs.LG2024-10

用视频或图片条件引导,让文字生成动作模型学会踢足球等新动作。

Enhancing Motion Variation in Text-to-Motion Models via Pose and Video Conditioned Editing

  • 以视频/图像为条件,结合已有动作先验生成新动作。
  • 用户研究显示生成动作真实度接近常见动作(如走、跑)。
  • 适合想扩展动作库的创作者和动画师使用。

从文本描述生成人体姿态序列的文本到动作模型受到广泛关注。然而,由于数据稀缺,这些模型能生成的动作范围仍然有限。例如,当前模型无法生成用脚内侧踢足球的动作,因为训练数据仅包含武术类踢法。本文提出一种新方法,利用短视频或图像作为条件来修改现有基础动作。该方法将模型对踢动作的理解作为先验,而足球踢腿的视频或图像作为后验,从而生成目标动作。通过引入这些额外模态作为条件,本方法可生成训练集中不存在的动作,突破文本-动作数据集的限制。26名参与者参与的用户研究表明,该方法生成的未见动作在真实感上与文本-动作数据集中常见动作(如HumanML3D中的行走、跑步、下蹲、踢腿)相当。

原文摘要 · Abstract (English)

Text-to-motion models that generate sequences of human poses from textual descriptions are garnering significant attention. However, due to data scarcity, the range of motions these models can produce is still limited. For instance, current text-to-motion models cannot generate a motion of kicking a football with the instep of the foot, since the training data only includes martial arts kicks. We propose a novel method that uses short video clips or images as conditions to modify existing basic motions. In this approach, the model's understanding of a kick serves as the prior, while the video or image of a football kick acts as the posterior, enabling the generation of the desired motion. By incorporating these additional modalities as conditions, our method can create motions not present in the training set, overcoming the limitations of text-motion datasets. A user study with 26 participants demonstrated that our approach produces unseen motions with realism comparable to commonly represented motions in text-motion datasets (e.g., HumanML3D), such as walking, running, squatting, and kicking.

动作生成视频条件文本到动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。