arXiv:2409.15904cs.CV2024-09被引 41

统一动作生成与理解,支持多层级文本控制和细粒度动作解析

Unimotion: Unifying 3D Human Motion Synthesis and Understanding

  • 采用全局+局部文本联合控制,实现灵活动作生成
  • 在HumanML3D上达最优帧级文本到动作性能
  • 可自动生成动作描述,适合动画编辑与数据标注

我们提出Unimotion,首个能同时实现灵活动作控制与帧级动作理解的统一模型。现有方法或仅支持全局文本控制,或仅支持逐帧脚本控制,且无法输出与动作对应的帧级文本。Unimotion首次设计支持全局文本、局部帧级文本或两者结合的控制方式,并原生输出与生成姿态配对的局部文本,使用户可明确知晓每个时刻的动作内容,推动多种新应用:1)分层控制,支持不同粒度的动作指令;2)为已有动作捕捉数据或YouTube视频自动生成动作描述;3)支持通过文本编辑修改动作。Unimotion在标准HumanML3D数据集的帧级文本到动作任务上达到当前最优性能。预训练模型与代码已开源。

原文摘要 · Abstract (English)

We introduce Unimotion, the first unified multi-task human motion model capable of both flexible motion control and frame-level motion understanding. While existing works control avatar motion with global text conditioning, or with fine-grained per frame scripts, none can do both at once. In addition, none of the existing works can output frame-level text paired with the generated poses. In contrast, Unimotion allows to control motion with global text, or local frame-level text, or both at once, providing more flexible control for users. Importantly, Unimotion is the first model which by design outputs local text paired with the generated poses, allowing users to know what motion happens and when, which is necessary for a wide range of applications. We show Unimotion opens up new applications: 1.) Hierarchical control, allowing users to specify motion at different levels of detail, 2.) Obtaining motion text descriptions for existing MoCap data or YouTube videos 3.) Allowing for editability, generating motion from text, and editing the motion via text edits. Moreover, Unimotion attains state-of-the-art results for the frame-level text-to-motion task on the established HumanML3D dataset. The pre-trained model and code are available available on our project page at https://coral79.github.io/uni-motion/.

动作生成多任务学习文本控制动作理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。