用文本和单帧姿态生成逼真人体动作,支持多任务统一建模。
MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding
- 将文本与姿态转为离散符号,接入大语言模型统一生成
- 在3D全身动作生成上达到领先性能,手部动作细节更精细
- 适合数字人、动画制作等需要多模态动作控制的场景
从描述性文本生成逼真人体动作近年来受到广泛关注,尤其在数字人需求推动下。尽管进展显著,现有方法常受限于控制模态少、任务专一、仅关注身体运动表示。本文提出MotionGPT-2,一种统一的大规模运动-语言模型(LMLM),解决上述问题。该模型通过预训练大语言模型(LLM)支持多种运动相关任务,并将文本与单帧姿态等多模态输入量化为离散、可被LLM理解的令牌,融入其词汇体系。这些令牌构成统一提示,通过预训练-微调范式引导生成运动输出。我们还展示了创新的运动离散化框架Part-Aware VQVAE,使MotionGPT-2在挑战性的3D全身动作生成任务中表现出高度适应性,确保身体与手部动作的细粒度表征。大量实验与可视化验证了方法的有效性,在动作生成、动作描述、泛化动作补全任务中均表现优异。
原文摘要 · Abstract (English)
Generating lifelike human motions from descriptive texts has experienced remarkable research focus in the recent years, propelled by the emerging requirements of digital humans.Despite impressive advances, existing approaches are often constrained by limited control modalities, task specificity, and focus solely on body motion representations.In this paper, we present MotionGPT-2, a unified Large Motion-Language Model (LMLM) that addresses these limitations. MotionGPT-2 accommodates multiple motion-relevant tasks and supporting multimodal control conditions through pre-trained Large Language Models (LLMs). It quantizes multimodal inputs-such as text and single-frame poses-into discrete, LLM-interpretable tokens, seamlessly integrating them into the LLM's vocabulary. These tokens are then organized into unified prompts, guiding the LLM to generate motion outputs through a pretraining-then-finetuning paradigm. We also show that the proposed MotionGPT-2 is highly adaptable to the challenging 3D holistic motion generation task, enabled by the innovative motion discretization framework, Part-Aware VQVAE, which ensures fine-grained representations of body and hand movements. Extensive experiments and visualizations validate the effectiveness of our method, demonstrating the adaptability of MotionGPT-2 across motion generation, motion captioning, and generalized motion completion tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。