用自然语言生成动作,让角色动起来更直观。
Text-driven Motion Generation: Overview, Challenges and Directions
- 分架构和动作表示两类方法,梳理主流技术路线
- 总结常用数据集与评估方式,明确领域现状
- 指出现有挑战与未来方向,适合研究者参考
文本驱动的动作生成为直接从自然语言生成人类动作提供了一种强大且直观的方法。该方法无需预定义动作输入,使动画角色控制更加灵活易用,适用于虚拟现实、游戏、人机交互及机器人等领域。本文首先回顾传统动作合成范式,即基于初始序列预测未来姿态,常以动作标签为条件;随后系统梳理现代文本到动作生成方法,从架构角度分为基于变分自编码器(VAE)、扩散模型及混合模型;从动作表示角度区分离散与连续生成策略。此外,还探讨了当前广泛使用的数据集、评估方法与最新基准,全面呈现该领域的进展。本文旨在厘清当前研究态势,揭示关键挑战与局限,并提出未来探索方向,期望为语言驱动的人体动作合成研究提供有价值的起点。
原文摘要 · Abstract (English)
Text-driven motion generation offers a powerful and intuitive way to create human movements directly from natural language. By removing the need for predefined motion inputs, it provides a flexible and accessible approach to controlling animated characters. This makes it especially useful in areas like virtual reality, gaming, human-computer interaction, and robotics. In this review, we first revisit the traditional perspective on motion synthesis, where models focused on predicting future poses from observed initial sequences, often conditioned on action labels. We then provide a comprehensive and structured survey of modern text-to-motion generation approaches, categorizing them from two complementary perspectives: (i) architectural, dividing methods into VAE-based, diffusion-based, and hybrid models; and (ii) motion representation, distinguishing between discrete and continuous motion generation strategies. In addition, we explore the most widely used datasets, evaluation methods, and recent benchmarks that have shaped progress in this area. With this survey, we aim to capture where the field currently stands, bring attention to its key challenges and limitations, and highlight promising directions for future exploration. We hope this work offers a valuable starting point for researchers and practitioners working to push the boundaries of language-driven human motion synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。