MotionGlot可跨身体形态生成动作,提升35.3%性能。
MotionGlot: A Multi-Embodied Motion Generation Model
- 借鉴大模型训练思路,设计动作生成指令模板。
- 6项任务平均性能提升35.3%,支持多形态动作生成。
- 适合机器人、动画等领域研究者使用。
本文提出MotionGlot,一种可在不同身体形态(如四足机器人和人形)上生成动作的模型,具备不同动作维度的适应能力。通过借鉴大语言模型(LLMs)的成熟训练流程,我们设计了专用于运动生成任务的指令微调模板,验证了该范式在多形态运动生成任务中的有效性。我们在6个任务上展示了MotionGlot的能力,平均性能提升35.3%。此外,我们构建了两个新数据集:(1) 约48,000条由专家控制的四足动物运动轨迹,配以基于方向的文本标注;(2) 超过23,000条用于人形动作生成的情境文本提示。最后,通过硬件实验验证了系统在真实场景中的可行性。
原文摘要 · Abstract (English)
This paper introduces MotionGlot, a model that can generate motion across multiple embodiments with different action dimensions, such as quadruped robots and human bodies. By leveraging the well-established training procedures commonly used in large language models (LLMs), we introduce an instruction-tuning template specifically designed for motionrelated tasks. Our approach demonstrates that the principles underlying LLM training can be successfully adapted to learn a wide range of motion generation tasks across multiple embodiments with different action dimensions. We demonstrate the various abilities of MotionGlot on a set of 6 tasks and report an average improvement of 35.3% across tasks. Additionally, we contribute two new datasets: (1) a dataset of expert-controlled quadruped locomotion with approximately 48,000 trajectories paired with direction-based text annotations, and (2) a dataset of over 23,000 situational text prompts for human motion generation tasks. Finally, we conduct hardware experiments to validate the capabilities of our system in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。