让动作生成支持风格、文本、轨迹的精细控制,用户可自由组合指定动作特征。
ACMo: Attribute Controllable Motion Generation
- 分离属性条件,通过扩散模型实现文本与动作解耦学习。
- 引入运动适配器,快速微调未见动作模式,支持多模态生成。
- 结合大语言模型规划器,实现用户友好交互,提升可控性与泛化能力。
风格、细粒度文本和轨迹等属性是描述动作的重要条件。然而,现有方法在动作属性的精准控制和对未见动作的泛化能力方面存在不足。本文提出一种属性可控的动作生成架构(ACMo),通过解耦各类条件并分别控制来解决上述问题。首先,采用属性扩散模型,通过解耦文本与动作学习,提升文本到动作的生成性能,依赖预训练模型实现可控生成。其次,引入运动适配器,可快速微调未见的动作模式,其运动提示输入支持多模态文本到动作生成,能捕捉用户指定的风格。最后,提出大语言模型规划器,利用局部知识桥接未见属性与数据集中特定文本,实现用户友好的交互。本方法引入运动提示机制,实现风格化生成,支持细粒度且友好的属性控制,性能达到当前最优水平。
原文摘要 · Abstract (English)
Attributes such as style, fine-grained text, and trajectory are specific conditions for describing motion. However, existing methods often lack precise user control over motion attributes and suffer from limited generalizability to unseen motions. This work introduces an Attribute Controllable Motion generation architecture, to address these challenges via decouple any conditions and control them separately. Firstly, we explored the Attribute Diffusion Model to imporve text-to-motion performance via decouple text and motion learning, as the controllable model relies heavily on the pre-trained model. Then, we introduce Motion Adpater to quickly finetune previously unseen motion patterns. Its motion prompts inputs achieve multimodal text-to-motion generation that captures user-specified styles. Finally, we propose a LLM Planner to bridge the gap between unseen attributes and dataset-specific texts via local knowledage for user-friendly interaction. Our approach introduces the capability for motion prompts for stylize generation, enabling fine-grained and user-friendly attribute control while providing performance comparable to state-of-the-art methods. Project page: https://mjwei3d.github.io/ACMo/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。