arXiv:2501.18726cs.CV2025-01被引 3

提出高效可控的3D动作生成方法,支持实时应用与精细关节控制。

Strong and Controllable 3D Motion Generation

  • 采用定制化闪速注意力与一致性模型,加速生成过程
  • 在HumanML3D数据集上实现10倍以上推理速度提升
  • 引入Motion ControlNet,实现更精准的关节级动作控制

人体动作生成在影视制作、游戏开发、增强现实/虚拟现实及人机交互等领域具有广泛应用。现有方法主要依赖基于扩散或自回归的文本到动作生成模型,但面临两大挑战:一是生成过程耗时,难以满足游戏、机器人操作等实时场景需求;二是通常学习相对动作表示,难以实现精确的关节级控制。为此,本文提出一种简洁高效的架构,包含两个核心组件:首先,通过定制闪速线性注意力,优化基于Transformer的扩散模型在人体动作生成中的硬件效率与计算复杂度;进一步,在动作隐空间中定制一致性模型以加速生成。其次,提出Motion ControlNet,相比以往方法显著提升关节级动作控制精度。该工作推动文本到动作生成技术向真实应用迈进。

原文摘要 · Abstract (English)

Human motion generation is a significant pursuit in generative computer vision with widespread applications in film-making, video games, AR/VR, and human-robot interaction. Current methods mainly utilize either diffusion-based generative models or autoregressive models for text-to-motion generation. However, they face two significant challenges: (1) The generation process is time-consuming, posing a major obstacle for real-time applications such as gaming, robot manipulation, and other online settings. (2) These methods typically learn a relative motion representation guided by text, making it difficult to generate motion sequences with precise joint-level control. These challenges significantly hinder progress and limit the real-world application of human motion generation techniques. To address this gap, we propose a simple yet effective architecture consisting of two key components. Firstly, we aim to improve hardware efficiency and computational complexity in transformer-based diffusion models for human motion generation. By customizing flash linear attention, we can optimize these models specifically for generating human motion efficiently. Furthermore, we will customize the consistency model in the motion latent space to further accelerate motion generation. Secondly, we introduce Motion ControlNet, which enables more precise joint-level control of human motion compared to previous text-to-motion generation methods. These contributions represent a significant advancement for text-to-motion generation, bringing it closer to real-world applications.

动作生成扩散模型控制实时

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。