arXiv:2506.03191cs.CVcs.AI2025-06被引 9

用文本生成逼真人体动作,大模型让指令与动作更对得上。

Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward

  • 用文本控制动作生成,让语言指令精准引导运动序列。
  • 结合LLM提升语义理解,使动作更连贯、符合上下文。
  • 适合游戏、动画、机器人等领域,推动智能动作合成发展。

本文深入调研了多模态生成人工智能与自回归大语言模型在人体运动理解与生成中的应用,探讨了新兴方法与架构,揭示其在实现逼真、多样运动合成方面的潜力。聚焦文本与运动模态,研究如何利用文本描述指导复杂类人运动序列的生成。分析了自回归模型、扩散模型、生成对抗网络(GANs)、变分自编码器(VAEs)及基于Transformer的模型在运动质量、计算效率和适应性方面的优劣。重点阐述了文本条件运动生成的最新进展,即通过文本输入精确控制与优化运动输出。引入大语言模型进一步增强语义对齐能力,提升指令与动作的一致性与上下文相关性。系统性综述强调了文本到运动生成与大模型架构在医疗、人形机器人、游戏、动画及辅助技术中的变革潜力,同时指出生成高效且真实人体运动仍面临挑战。

原文摘要 · Abstract (English)

This paper presents an in-depth survey on the use of multimodal Generative Artificial Intelligence (GenAI) and autoregressive Large Language Models (LLMs) for human motion understanding and generation, offering insights into emerging methods, architectures, and their potential to advance realistic and versatile motion synthesis. Focusing exclusively on text and motion modalities, this research investigates how textual descriptions can guide the generation of complex, human-like motion sequences. The paper explores various generative approaches, including autoregressive models, diffusion models, Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), and transformer-based models, by analyzing their strengths and limitations in terms of motion quality, computational efficiency, and adaptability. It highlights recent advances in text-conditioned motion generation, where textual inputs are used to control and refine motion outputs with greater precision. The integration of LLMs further enhances these models by enabling semantic alignment between instructions and motion, improving coherence and contextual relevance. This systematic survey underscores the transformative potential of text-to-motion GenAI and LLM architectures in applications such as healthcare, humanoids, gaming, animation, and assistive technologies, while addressing ongoing challenges in generating efficient and realistic human motion.

动作生成多模态大模型文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。